Technical field guide
Understand the word.
Keep following the idea.
Plain-language explanations connect the technical terms in these articles and stories. Follow related concepts, then open the engineering that puts them to work.
Hover over a dotted term for 5 seconds to fill its frame and lock the explanation. Click, tap or press Enter to lock immediately. Explanations close after 5 seconds away; hovering a child tooltip or keeping keyboard focus within the explanation keeps its parents open.
404 explanations
Cloud, networking and security
ACL
A list of access rules applied at a particular boundary. Different services interpret rules differently, so an ACL should be evaluated as part of the complete access path.
Connected ideas
Follow the engineering
AI inference and model serving
Activation
An intermediate value produced while a model executes. Activations and temporary workspaces add to the memory boundary beyond the model’s stored weights.
Connected ideas
Follow the engineering
Hardware and execution
Active lanes
The lanes participating in a particular instruction. This helps explain control flow and masking, but it is a local property, not a prediction of overall GPU throughput.
Connected ideas
Follow the engineering
AI inference and model serving
Admission control
A decision about whether work is allowed to enter a system at that moment. Capacity-aware admission can protect latency and memory, but rejected requests must remain visible in performance reporting.
Connected ideas
Follow the engineering
Kubernetes and production operations
Affinity
Placement rules or preferences expressing which resources should be together or apart. Hard constraints, preferences and failure-domain spreading have different effects on schedulability and resilience.
Connected ideas
Follow the engineering
AI inference and model serving
AI
A broad category of systems that perform tasks such as prediction, generation or decision support. The label does not establish reliability or authority; a useful deployment still needs a bounded task, evaluation and operational controls.
Connected ideas
Follow the engineering
Cloud, networking and security
AKS
Azure’s managed Kubernetes service. Platform migration still needs explicit contracts for identity, storage, networking and model or application releases.
Connected ideas
Follow the engineering
Hardware and execution
ALU
An execution unit for supported arithmetic and logic operations. A GPU contains different kinds of pipelines, and the existence of many ALUs does not mean every instruction can use all of them at once.
Connected ideas
Follow the engineering
Hardware and execution
AND
A Boolean operation that is true only when both input bits are true. In a full adder, AND paths help determine when input bits generate or propagate a carry.
Connected ideas
Follow the engineering
Software, data and observability
API
A defined interface through which software requests data or actions from another component. The contract includes inputs, outputs, permissions and failure behavior, not just a URL that returns a response.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
API gateway
An entry point that applies selected routing and policy to API traffic. Gateway validation is one boundary; downstream services still need the authority and input checks appropriate to their actions.
Connected ideas
Follow the engineering
Cloud, networking and security
AppArmor
A Linux security mechanism for applying mandatory access policies to processes. Its effectiveness depends on the profile and enforcement mode, not merely on the feature being installed.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Approval gate
A required decision before a bounded action may execute. Approval should identify the actual action and scope; a general conversation about a task is not automatically permission for every external effect.
Connected ideas
Follow the engineering
Kernels and numerical computing
Arithmetic intensity
The ratio of arithmetic work to data movement at a specified memory boundary. Changing reuse or fusion can change this ratio, but the units and accounting boundary must remain explicit.
Connected ideas
Follow the engineering
Business value and cost
ARR
A recurring-revenue metric normalized to a year under a business’s definition. It is not equivalent to cash collected, profit, daily active users or the revenue a capacity model could support.
Connected ideas
Follow the engineering
Delivery and engineering tools
AST
A structured representation of source code after parsing. Syntax-aware tools can distinguish real code constructs from coincidental text, but semantic compatibility may require further analysis.
Connected ideas
Follow the engineering
Kernels and numerical computing
Asynchronous
Work whose submission and completion occur at different times. A host call returning does not necessarily mean GPU work or an external effect has completed.
Connected ideas
Follow the engineering
AI inference and model serving
Attention
A model operation that combines information from selected parts of a sequence. The exact attention layout and implementation affect compute, cache requirements and supported optimizations.
Connected ideas
Follow the engineering
Cloud, networking and security
Authentication
Establishing which identity is presenting a request. Authentication is separate from deciding whether that identity may perform the requested action.
Connected ideas
Follow the engineering
Cloud, networking and security
Authentication token
A value used to present access authority under a system’s rules. Scope, lifetime, storage and revocation matter; an authentication token must not be confused with the text tokens processed by an LLM.
Connected ideas
Follow the engineering
Cloud, networking and security
Authorization
Deciding whether a specific identity may perform a specific action on a resource. Model intent, successful login and execution isolation are not replacements for this decision.
Connected ideas
Follow the engineering
Kubernetes and production operations
Autoscaling
Changing allocated capacity in response to declared signals and policies. Pod scaling, node provisioning and application readiness happen at different layers and on different timelines.
Connected ideas
Follow the engineering
Software, data and observability
Availability zone
A provider-defined placement boundary intended to separate selected infrastructure failures. Cross-zone topology can also change latency and cost, and it is not a universal guarantee against correlated failure.
Connected ideas
Follow the engineering
Cloud, networking and security
AWS
A cloud provider offering infrastructure and managed services. Compare an exact workload and service boundary; the provider name alone does not establish equivalent cost, identity behavior or accelerator capacity.
Connected ideas
Follow the engineering
Cloud, networking and security
Azure
Microsoft’s cloud platform. Azure services and operating contracts need to be compared individually with alternatives, including data, identity and recovery behavior.
Connected ideas
Follow the engineering
Hardware and execution
B200
A Blackwell-generation NVIDIA accelerator used in several server offerings. Published device memory, system memory and allocatable runtime memory describe different boundaries, so a B200 label alone is not a serving-capacity guarantee.
Connected ideas
Follow the engineering
Business value and cost
B2B
A business model in which customers are other organizations. Technical capacity, procurement, adoption and revenue remain different parts of the commercial story.
Connected ideas
Follow the engineering
Hardware and execution
B300
A Blackwell Ultra-generation NVIDIA accelerator offered in specific server configurations. Compare the provider’s exact memory, networking and host specifications rather than mixing figures from different B300 systems.
Connected ideas
Follow the engineering
Software, data and observability
Backpressure
A mechanism that slows or rejects new work when downstream capacity is constrained. It protects a system from unbounded queues but must be visible to callers and performance reporting.
Connected ideas
Follow the engineering
Software, data and observability
Backup
A retained copy intended to support recovery from a defined failure or mistake. A backup policy is incomplete without testing restoration and identifying the data and dependencies it must recover.
Connected ideas
Follow the engineering
Hardware and execution
Bandwidth
The rate at which a resource can move data, such as bytes per second through memory or a network link. Application throughput can be lower because of access patterns, contention, protocol overhead or another bottleneck.
Connected ideas
Follow the engineering
Kernels and numerical computing
Bank conflict
Contention caused by particular access patterns to shared-memory banks. Padding can help some patterns, but should follow the actual lane-to-address mapping rather than being added as a ritual.
Connected ideas
Follow the engineering
Kernels and numerical computing
Barrier
A synchronization point where the participating work must meet a defined condition before proceeding. Incorrect participation can deadlock a kernel or expose data before it is ready.
Connected ideas
Follow the engineering
AI inference and model serving
Batching
Processing compatible work together to use resources more effectively. Larger batches can improve utilization but also change memory demand, queueing and per-request latency.
Connected ideas
Follow the engineering
Kernels and numerical computing
Benchmarking
A controlled comparison of behavior under a declared workload and measurement method. Useful benchmarks preserve conditions, failures and reproducibility instead of presenting an isolated peak as a general guarantee.
Connected ideas
Follow the engineering
Kernels and numerical computing
BF16
A 16-bit floating-point format with exponent range similar to Float32 but fewer fraction bits. Its behavior is not interchangeable with FP16; evaluate the application’s numerical and quality requirements.
Connected ideas
Follow the engineering
Hardware and execution
Blackwell
An NVIDIA accelerator architecture family that includes multiple device and system configurations. Architectural capabilities are useful context, but deployment and performance claims must identify the actual product and workload.
Connected ideas
Follow the engineering
Hardware and execution
Bottleneck
The resource or dependency currently limiting useful progress. Improving a different part of the system may have little effect; profiling and controlled changes help identify the active constraint.
Connected ideas
Follow the engineering
Kernels and numerical computing
Boundary mask
A condition that prevents work outside the valid data range. Masks must protect both reads and writes while preserving required synchronization for participating threads.
Connected ideas
Follow the engineering
Business value and cost
Break-even
The point where the compared costs or benefits balance under explicit assumptions. A break-even model needs comparable workload, time and operating boundaries rather than mixing unrelated pricing units.
Connected ideas
Follow the engineering
Delivery and engineering tools
C++
A systems programming language used for many performance-sensitive applications and CUDA kernels. Memory ownership, synchronization and supported numerical behavior remain explicit correctness responsibilities.
Connected ideas
Follow the engineering
Hardware and execution
Cache
Storage that keeps reusable data closer to the work that needs it. A cache can avoid repeated transfers, but its capacity, invalidation rules and hit rate determine whether that benefit appears in the real workload.
Connected ideas
Follow the engineering
Hardware and execution
Cache hit
An access satisfied by an existing cache entry. A useful measurement distinguishes which cache supplied the data and how often misses still require a more expensive path.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Capability
In security, a bounded ability to perform an operation on a resource. Possessing one capability should not imply unrelated authority, and capability descriptions must correspond to real enforcement.
Connected ideas
Follow the engineering
Business value and cost
CapEx
Spending accounted for as acquiring longer-lived assets, under applicable accounting rules. Hardware ownership comparisons should still include financing, useful lifetime, utilization and the operating costs around the asset.
Connected ideas
Follow the engineering
Software, data and observability
Cardinality
The number of distinct combinations represented by telemetry dimensions. Unbounded labels such as request IDs can multiply storage and query costs without making aggregate metrics more useful.
Connected ideas
Follow the engineering
Business value and cost
Cash savings
A reduction in actual relevant spending under a stated comparison. Staff time made available, faster delivery and modeled revenue timing are different benefits and should not be reported as the same cash saving.
Connected ideas
Follow the engineering
Software, data and observability
CDC
Capturing data changes so that another system can observe or apply them. Ordering, duplicate handling and recovery positions determine whether a migration or downstream view remains correct.
Connected ideas
Follow the engineering
Cloud, networking and security
cgroups
Linux mechanisms for organizing and controlling resource usage. Resource enforcement is distinct from isolating a process’s filesystem view or authorizing its external actions.
Connected ideas
Follow the engineering
Software, data and observability
Checkpoint
A recorded position or state from which work can resume under a defined recovery contract. A checkpoint is useful only when its related data and external effects can be reconciled correctly.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Checkpointing
Recording recoverable state at defined points in a computation or workflow. The checkpoint’s consistency with artifacts and external effects matters more than merely writing a file.
Connected ideas
Follow the engineering
Delivery and engineering tools
CI/CD
A shorthand for continuous integration and delivery or deployment practices. The exact pipeline must still say which steps validate, which mutate production, and where explicit approval is required.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Circuit breaker
A mechanism that limits calls to a failing or overloaded dependency under a recovery policy. It can reduce cascading load but must expose failure and recovery behavior to the rest of the application.
Connected ideas
Follow the engineering
Software, data and observability
CLI
A text-command interface for operating a program or system. Commands still have authority and side effects, so automation must constrain arguments, permissions and failure handling.
Connected ideas
Follow the engineering
Cloud, networking and security
Cloud computing
Using remotely operated compute, storage and managed services under provider contracts. A cloud decision includes identity, data movement, reliability, operating work and exit costs—not only a headline instance price.
Connected ideas
Follow the engineering
Cloud, networking and security
Cloud migration
Moving a workload and its operating contracts between infrastructure environments. A safe migration includes identity, state, validation and a reversible cutover, not just copying a deployment manifest.
Connected ideas
Follow the engineering
Kubernetes and production operations
Cluster
A group of machines or services managed as a cooperating system. Cluster membership alone does not establish uniform capacity, one failure domain or a shared memory allocation.
Connected ideas
Follow the engineering
Kubernetes and production operations
Cluster Autoscaler
Automation that changes available node capacity under its supported rules. Provisioned machines still need to become usable, and reducing desired Pods does not guarantee an equivalent reduction in billed nodes.
Connected ideas
Follow the engineering
Hardware and execution
CMOS
A common transistor technology for digital circuits, using complementary device behavior. A logic value represents a voltage range, and switching involves physical capacitance, power and delay.
Connected ideas
Follow the engineering
Kernels and numerical computing
Coalescing
Combining compatible memory accesses into efficient transactions. Adjacent-looking source code is not enough: derive which addresses the participating lanes actually access.
Connected ideas
Follow the engineering
AI inference and model serving
Cold start
The delay before a newly started service can handle useful work. For model serving this can include provisioning, image and artifact loading, initialization and warmup—not merely starting a process.
Connected ideas
Follow the engineering
AI inference and model serving
Collective
A coordinated communication operation involving a group of workers, such as reducing or distributing values. A slow participant or transport path can constrain the entire operation.
Connected ideas
Follow the engineering
Cloud, networking and security
Colocation
Placing owned or contracted equipment in a facility that supplies defined services such as space, power and connectivity. Equipment lifecycle and workload operations remain separate responsibilities unless the contract explicitly includes them.
Connected ideas
Follow the engineering
Kernels and numerical computing
Compilation
Translation and optimization of a program into another representation or executable form. A source-level change may not produce the expected device instructions; inspect and measure the generated result.
Connected ideas
Follow the engineering
Kernels and numerical computing
Compute-bound
A workload currently constrained by arithmetic or instruction execution rather than the selected memory path. Peak vendor compute rates are only relevant under the formats and operation types they describe.
Connected ideas
Follow the engineering
AI inference and model serving
Concurrency
The number of operations or requests overlapping at a given time. It is distinct from daily users and request rate, and it changes the amount of memory and capacity a service may need.
Connected ideas
Follow the engineering
Delivery and engineering tools
Configuration
Settings that influence how a system behaves. Configuration should have a clear source of truth and change process so that operators can distinguish intended changes from drift.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
CONNECT
An HTTP method used to establish a tunnel through a proxy under its policy. A tool’s network restrictions must account for such transport paths instead of validating only an obvious URL hostname.
Connected ideas
Follow the engineering
Software, data and observability
Consistency
The rules for what readers can observe as a system changes. Eventual convergence differs from immediately observing one agreed value, and the application must tolerate the chosen model.
Connected ideas
Follow the engineering
Kubernetes and production operations
Consolidation
Reducing or reshaping capacity so that useful work fits on fewer or more suitable resources. It must account for placement, disruption and readiness rather than simply deleting the least busy machine.
Connected ideas
Follow the engineering
Kubernetes and production operations
Container
A packaged application execution environment that shares underlying operating-system facilities. A container is not automatically a complete security sandbox or an independent virtual machine.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Container image
A packaged filesystem and metadata used to create container execution. An image should be identifiable and reviewed, but runtime configuration, secrets, data and platform policy remain separate inputs.
Connected ideas
Follow the engineering
Kubernetes and production operations
Container runtime
Software that creates and manages container execution using supported operating-system facilities. Runtime configuration is one layer of the workload’s security and device-access contract.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Content digest
A deterministic fingerprint of content produced by a hashing algorithm. A digest can bind a review or release to exact bytes, but authenticity still depends on who is trusted to approve or publish that digest.
Connected ideas
Follow the engineering
AI inference and model serving
Context window
The input and retained sequence length a model or service can support under its configuration. Longer contexts affect work and cache requirements; a supported maximum is not a promise of the same latency at every length.
Connected ideas
Follow the engineering
AI inference and model serving
Continuous batching
A serving technique that admits and retires requests as their generation progresses rather than holding one fixed batch until every request finishes. Scheduling policy still determines fairness, cache pressure and latency.
Connected ideas
Follow the engineering
Delivery and engineering tools
Continuous delivery
Practices that keep a change ready for a controlled release. Build, review, publication and deployment remain distinct actions with their own evidence and authorization boundaries.
Connected ideas
Follow the engineering
Delivery and engineering tools
Continuous integration
Automated checks and integration practices for source changes. Passing CI provides evidence about the checks that ran, not permission to publish or a guarantee of production behavior.
Connected ideas
Follow the engineering
Business value and cost
Contribution margin
Revenue remaining after the variable costs included in a defined accounting boundary. It differs from gross revenue, cash savings and profit, and must not be added to overlapping benefits in an ROI story.
Connected ideas
Follow the engineering
Kubernetes and production operations
Control plane
The components that decide and reconcile system configuration, rather than directly serving every application request. Control-plane health and application data-path health are separate operational concerns.
Connected ideas
Follow the engineering
Kubernetes and production operations
Controller
A component that repeatedly compares observed state with a declared target and attempts to close the gap. Clear ownership matters: two controllers trying to change the same property can fight each other.
Connected ideas
Follow the engineering
Kernels and numerical computing
Correctness
Whether an implementation satisfies its stated contract for supported inputs and failure cases. Performance comparisons are meaningful only after the result and boundaries have been checked independently.
Connected ideas
Follow the engineering
Hardware and execution
CPU
A general-purpose processor that runs operating-system and application instructions. In GPU systems, CPU work such as request handling, tokenization and batch preparation can still limit the service.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Credential
Information used to establish identity or access, such as a key, password or authentication token. Credentials need controlled scope and custody and should not be embedded in logs, examples or broadly shared application state.
Connected ideas
Follow the engineering
Kubernetes and production operations
CSI
An interface used to integrate storage systems with container orchestration. A CSI integration exposes capabilities; it does not by itself establish data durability or a tested recovery procedure.
Connected ideas
Follow the engineering
Kernels and numerical computing
CUDA
NVIDIA’s programming platform and execution model for supported GPUs. CUDA code must define work ownership, memory access and synchronization; using CUDA does not automatically improve an application over a tuned library.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
CUDA graph
A recorded execution structure that can reduce some submission overhead for supported work. Capture conditions, memory use and workload compatibility remain part of a valid inference comparison.
Connected ideas
Follow the engineering
Kernels and numerical computing
CUDA grid
The collection of thread blocks launched for a CUDA kernel. Grid dimensions describe how work is divided, not a guarantee that every block runs simultaneously.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
CUDA runtime
Libraries and APIs that help applications manage supported CUDA execution and resources. The runtime is separate from the operating-system driver and from Kubernetes’ workload placement decisions.
Connected ideas
Follow the engineering
Kubernetes and production operations
DaemonSet
A Kubernetes controller used to run a workload on eligible nodes, often for node-level services. Its placement and rollout behavior differs from a normal Deployment replica count.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Data lineage
The recorded relationship between data, transformations and resulting artifacts. Lineage helps explain and reproduce a result, but it must include the dependencies that actually affected the work.
Connected ideas
Follow the engineering
Software, data and observability
Data migration
Changing the location, shape or ownership of data. Rollback and reconciliation must cover writes made during the transition, not only the deployment manifests.
Connected ideas
Follow the engineering
Software, data and observability
Data ownership
Responsibility for the meaning, changes and access rules of a data set. Clear ownership is essential when extracting services, so two deployments do not accidentally become competing authorities for one record.
Connected ideas
Follow the engineering
Kubernetes and production operations
Data plane
The path where application requests or data actually move and are processed. A healthy configuration API does not prove that this path meets user-facing service requirements.
Connected ideas
Follow the engineering
Software, data and observability
Data retention
Rules for how long data is kept and when it is removed. Retention should match purpose and privacy requirements, rather than growing indefinitely because storage is available.
Connected ideas
Follow the engineering
Hardware and execution
Data transfer
Movement of data between memory, devices or services. Transfers can dominate elapsed time even when computation is fast, so include the relevant movement in application-level measurements.
Connected ideas
Follow the engineering
Software, data and observability
Database
A system for storing and retrieving data under defined consistency and durability guarantees. Choosing a database is also choosing its failure, recovery and operating boundaries.
Connected ideas
Follow the engineering
Software, data and observability
Deadline
A bound on how long a caller will wait or allow work to continue. A timeout does not prove that the underlying operation did nothing, so recovery needs to account for uncertain completion.
Connected ideas
Follow the engineering
AI inference and model serving
Decode
The phase that produces additional output tokens from the current sequence state. Decode performance depends on cache traffic, batch composition and scheduling, not just the speed of prefill.
Connected ideas
Follow the engineering
Delivery and engineering tools
Dependency
Something a component relies on to build, run or provide its promised behavior. Hidden or mutable dependencies make reproduction, security review and recovery harder.
Connected ideas
Follow the engineering
Kubernetes and production operations
Deployment
A Kubernetes controller commonly used for replaceable application replicas and rolling updates. Its rollout and availability settings interact with the application’s readiness and shutdown behavior.
Connected ideas
Follow the engineering
Kubernetes and production operations
Desired state
The state an operator or controller asks a system to reach. It is distinct from observed, scheduled and ready capacity, especially while resources are starting or failing.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Deterministic
Producing the same result from the same relevant inputs under a defined execution contract. Hidden state, timing, environment or random choices can break that assumption if they are left outside the input boundary.
Connected ideas
Follow the engineering
Kubernetes and production operations
Device plugin
A Kubernetes extension that advertises and allocates supported device resources to workloads. Device allocation is not the same as model capacity, shared-memory safety or useful inference throughput.
Connected ideas
Follow the engineering
Hardware and execution
DGX
An NVIDIA integrated system product family for accelerated computing. A DGX specification describes a whole system; individual GPU memory and host or aggregate system resources must still be separated.
Connected ideas
Follow the engineering
AI inference and model serving
Distributed inference
Serving a model using multiple devices or processes. Partitioning, communication and scheduling create additional constraints; aggregate compute and memory figures are not a single-device performance result.
Connected ideas
Follow the engineering
Software, data and observability
Distributed system
A system whose components coordinate across process or network boundaries. Partial failures, delay and uncertain outcomes make its contract different from a single in-process function call.
Connected ideas
Follow the engineering
Software, data and observability
Distributed tracing
Following an operation across process and service boundaries using related spans. Tracing helps explain a request path, but sampling and missing propagation can limit what the trace proves.
Connected ideas
Follow the engineering
Hardware and execution
Divergence
A situation where lanes in a group take different control-flow paths. The hardware must execute the required paths under masks, which can reduce useful lane participation.
Connected ideas
Follow the engineering
Hardware and execution
DMA
A mechanism allowing supported devices to move data without the CPU copying each byte through its normal instruction path. Authorization, mappings, synchronization and transfer cost still need to be handled.
Connected ideas
Follow the engineering
Cloud, networking and security
DNS
The system used to resolve names into records such as network addresses. Name resolution, network reachability and application authorization are separate checks.
Connected ideas
Follow the engineering
Delivery and engineering tools
Docker
Tooling and formats commonly associated with packaging applications into container images. An image captures part of a deployment, not all runtime configuration, data or security policy.
Connected ideas
Follow the engineering
Hardware and execution
DRAM
Memory that stores bits using cells that require periodic refresh. DRAM offers much more capacity than registers, but an application must account for memory access latency and bandwidth.
Connected ideas
Follow the engineering
Software, data and observability
Durability
The guarantee that acknowledged data survives specified failures. Durability, availability and backup coverage describe different properties and need separate evidence.
Connected ideas
Follow the engineering
Kernels and numerical computing
E2M1
A low-bit floating-point encoding with a particular exponent and fraction allocation. The encoding alone does not account for the scale metadata or accumulation behavior of a complete quantized operation.
Connected ideas
Follow the engineering
Kernels and numerical computing
E4M3
A floating-point format designation describing exponent and fraction-bit allocation. It identifies a numerical representation, not a guarantee of acceptable model quality or a particular kernel implementation.
Connected ideas
Follow the engineering
Cloud, networking and security
EC2
AWS’s virtual compute service with multiple instance families and configurations. Instance CPU, memory, network and storage figures describe a specific offering and must be kept within that boundary.
Connected ideas
Follow the engineering
Cloud, networking and security
EFA
AWS networking functionality designed for supported high-performance distributed workloads. Its transport and configuration are distinct from local NVLink connections and ordinary application-network traffic.
Connected ideas
Follow the engineering
Cloud, networking and security
Egress
Traffic leaving a defined system or provider boundary. Egress can affect both security and cost, and the relevant boundary must be specified before comparing provider quotes.
Connected ideas
Follow the engineering
Cloud, networking and security
EKS
AWS’s managed Kubernetes service. Managed control-plane functions do not remove application, node, identity, cost or workload-recovery responsibilities.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Embedding
A numerical representation used to compare or process items such as text. Similarity in embedding space is a modeling choice, not proof of factual correctness or permission to expose a retrieved item.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Encryption
Transforming data so that access depends on the relevant key and protocol. Encryption does not replace authorization, careful key custody or decisions about what data should be retained.
Connected ideas
Follow the engineering
Software, data and observability
Error budget
The allowed amount of failure under a defined service objective and measurement window. It can guide reliability and change decisions, but only when the indicator reflects the service that users need.
Connected ideas
Follow the engineering
Delivery and engineering tools
etcd
A distributed key-value store used by Kubernetes for control-plane state. It is an important platform dependency with its own consistency, backup and recovery requirements.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Evaluation
Checking a model or workflow against defined tasks, quality requirements and failure cases. A plausible demonstration is not sufficient evidence for deploying a system that can affect users or external resources.
Connected ideas
Follow the engineering
Kubernetes and production operations
Eviction
A request or action that removes a workload instance from its current placement. The cause and API path matter: voluntary disruption policy does not cover every way a Pod can stop.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Experiment tracking
Recording which inputs, settings, code and results belonged to a model experiment. Tracking is useful only when the referenced artifacts and execution conditions remain identifiable.
Connected ideas
Follow the engineering
Software, data and observability
Failure domain
A boundary within which one incident can affect multiple components together. Spreading resources only improves resilience if the chosen domains correspond to real shared dependencies.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Feature store
A system for managing model input features under defined consistency and serving contracts. Training and serving need compatible meanings and transformations, not merely similarly named columns.
Connected ideas
Follow the engineering
Kernels and numerical computing
FFI
A boundary for calling code implemented in another language or runtime. Argument ownership, layouts, lifetimes and error handling remain part of the contract even when the call looks small.
Connected ideas
Follow the engineering
Business value and cost
FinOps
Practices for connecting cloud spending to useful business and engineering decisions. Cost reporting is a starting point; verified savings require a change that actually reduces relevant spend without losing the service objective.
Connected ideas
Follow the engineering
Kernels and numerical computing
Float16
A 16-bit floating-point format with a different range and precision tradeoff from Float32 or BF16. It can reduce memory use, but supported operations and acceptable numerical error must be checked.
Connected ideas
Follow the engineering
Kernels and numerical computing
Float32
A 32-bit floating-point format commonly used for numerical arrays and reference calculations. Its finite range and rounding still require care; it is not exact real-number arithmetic.
Connected ideas
Follow the engineering
Kernels and numerical computing
FLOP
A unit of floating-point work under a stated counting convention. FLOP/s measures a rate, while a FLOP count measures work; confusing the two breaks performance and roofline calculations.
Connected ideas
Follow the engineering
Delivery and engineering tools
Flux
GitOps controllers that reconcile supported declared state from versioned sources. Flux and infrastructure tools need disjoint ownership so they do not compete over the same fields or resources.
Connected ideas
Follow the engineering
Kernels and numerical computing
FMA
An operation that computes a multiplication and addition with a defined rounding behavior. FLOP conventions commonly count its multiply and add separately, so state the accounting used in performance figures.
Connected ideas
Follow the engineering
Kernels and numerical computing
FP4
A family of 4-bit floating-point representations. The small encoding makes scaling and numerical validation especially important; storage savings do not by themselves prove useful inference quality or speed.
Connected ideas
Follow the engineering
Kernels and numerical computing
FP8
A family of 8-bit floating-point formats, not one universal numerical contract. Format, scaling, accumulation and kernel support affect both memory use and model quality.
Connected ideas
Follow the engineering
Hardware and execution
Full adder
A circuit that adds two input bits and a carry-in, producing a sum bit and carry-out. The two output bits have different place values and together represent the complete result.
Connected ideas
Follow the engineering
Kernels and numerical computing
GB
A decimal unit: one gigabyte is one billion bytes. Capacity in GB and bandwidth in GB/s are different quantities; neither is interchangeable with binary GiB without conversion.
Connected ideas
Follow the engineering
Hardware and execution
GB200
An NVIDIA Grace Blackwell product designation used in specific combined CPU/GPU platforms. Its system boundaries differ from an individual B200 GPU, so numbers cannot be transferred between them without qualification.
Connected ideas
Follow the engineering
Cloud, networking and security
GCP
Google’s cloud platform. Migration planning needs to account for service behavior and data/identity boundaries, not assume that a similarly named service is interchangeable with another provider’s offering.
Connected ideas
Follow the engineering
Kernels and numerical computing
GEMM
A standard matrix-multiplication operation, often expressed with scaling and accumulation. Matrix shape, layout and numerical format affect which implementation is appropriate.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
GET
An HTTP method intended to retrieve a resource without requesting a state-changing operation. The application still needs correct authorization and implementation; the method label is not a security boundary by itself.
Connected ideas
Follow the engineering
Kernels and numerical computing
GiB
A binary unit equal to 2^30 bytes. Keep GiB distinct from decimal GB when reading device inventory, estimating model fit and converting bytes into memory capacity.
Connected ideas
Follow the engineering
Cloud, networking and security
GID
A group identifier within a system’s access-control model. Group membership and filesystem or resource permissions must be evaluated together.
Connected ideas
Follow the engineering
Delivery and engineering tools
Git
A version-control system for recording and combining changes. A Git commit identifies source history; it does not itself prove that a build passed or that the corresponding artifact was deployed.
Connected ideas
Follow the engineering
Delivery and engineering tools
GitOps
Managing declared system state through reviewed Git changes and reconciliation. Clear field ownership and drift handling matter; GitOps should not become a second uncontrolled writer over the same resources as another tool.
Connected ideas
Follow the engineering
Cloud, networking and security
GKE
Google Cloud’s managed Kubernetes service. Its operating modes and integrations have their own contracts; portability must be checked beyond the Kubernetes API objects.
Connected ideas
Follow the engineering
Delivery and engineering tools
Go
A programming language commonly used for services and infrastructure software. It is one implementation choice; correctness, observability and ownership boundaries matter more than the language label alone.
Connected ideas
Follow the engineering
AI inference and model serving
Goodput
The rate of successful work that satisfies the declared quality and service requirements. It avoids calling rejected, failed or late requests a performance win.
Connected ideas
Follow the engineering
Hardware and execution
GPU
A processor built to run many suitable operations in parallel. A GPU is useful when the workload exposes enough parallel work and can feed its execution units; a larger device does not automatically make an application faster.
Connected ideas
Follow the engineering
Cloud, networking and security
GPU hosting
Running workloads on rented or operated accelerator infrastructure. Compare delivered service capability, idle time, networking and operating responsibility rather than treating a GPU label as a complete offer.
Connected ideas
Follow the engineering
Kernels and numerical computing
GPU kernel
A function launched to execute work on a GPU. Its correctness and speed depend on the supported inputs, mapping of work to threads, resource use and the surrounding launch and transfer costs.
Connected ideas
Follow the engineering
Hardware and execution
GPU memory
Memory available to a GPU for weights, arrays, intermediate results and runtime state. Usable capacity depends on the actual device and allocations; adding several GPUs does not automatically create one shared allocation.
Connected ideas
Follow the engineering
Hardware and execution
GPU memory bandwidth
The rate at which data can move through a GPU memory interface. A kernel may achieve less than the published maximum because its transactions, cache behavior and instruction dependencies differ from the vendor test.
Connected ideas
Follow the engineering
Kubernetes and production operations
GPU Operator
Automation for managing supported GPU software components in Kubernetes, such as drivers, runtime integration and device services. It manages lifecycle plumbing, not the quality or capacity of a particular inference model.
Connected ideas
Follow the engineering
Cloud, networking and security
GPU passthrough
Assigning a physical GPU or device to a guest under a supported isolation setup. This differs from subdividing the device or time-sharing it, and recovery and access rules still need to be explicit.
Connected ideas
Follow the engineering
AI inference and model serving
GQA
An attention layout in which multiple query heads share groups of key/value heads. The actual key/value-head count matters when estimating cache storage; using the query-head count can give the wrong result.
Connected ideas
Follow the engineering
Kubernetes and production operations
Graceful shutdown
Stopping a service in a way that honors in-flight work and resource cleanup within a declared time budget. Routing changes and process termination need to agree about when new work stops arriving.
Connected ideas
Follow the engineering
Business value and cost
Gross revenue
Revenue before the costs and deductions specified by the accounting model. Infrastructure may enable capacity or delivery, but a scaling model does not prove that customer demand or revenue was created.
Connected ideas
Follow the engineering
Software, data and observability
gRPC
An RPC framework for defining and calling service operations across a network. Transport and schema support do not eliminate application-level deadlines, failure semantics or compatibility requirements.
Connected ideas
Follow the engineering
Hardware and execution
H100
An NVIDIA Hopper-generation accelerator used for training and inference. Comparisons with newer hardware need the same model, precision, request mix and service objectives, not only a peak-compute number.
Connected ideas
Follow the engineering
Hardware and execution
H200
An NVIDIA Hopper-generation accelerator with a different memory configuration from H100 offerings. Extra memory may change model fit and batching, but useful performance still depends on the complete serving workload.
Connected ideas
Follow the engineering
Hardware and execution
HBM
Stacked memory placed close to a compatible processor to provide high aggregate bandwidth. HBM capacity and achieved bandwidth remain separate limits, and published bandwidth is not an application benchmark.
Connected ideas
Follow the engineering
Hardware and execution
HBM3e
A generation of high-bandwidth memory used by some recent accelerators. The memory generation alone does not identify a GPU configuration, its usable capacity, or the bandwidth a workload will sustain.
Connected ideas
Follow the engineering
Delivery and engineering tools
Helm
Packaging and templating tools commonly used for Kubernetes resources. A chart’s output still needs review for ownership, permissions, workload behavior and supported upgrade paths.
Connected ideas
Follow the engineering
Hardware and execution
HGX
An NVIDIA platform design used by server manufacturers to integrate multiple accelerators. It is not the same product boundary as a complete DGX system or a cloud instance.
Connected ideas
Follow the engineering
Software, data and observability
Histogram
A representation of how observations are distributed across value ranges. Histograms can support latency analysis, but bucket boundaries and aggregation methods affect the conclusions you can draw.
Connected ideas
Follow the engineering
Kubernetes and production operations
HPA
A Kubernetes controller that adjusts a workload’s desired replica count from configured signals and policy. Its raw ratio is not a Ready-Pod count, and HPA does not itself provision the nodes those Pods may need.
Connected ideas
Follow the engineering
Cloud, networking and security
HTTP
An application protocol for requests and responses. A transport-level success status is only one part of a service’s contract; the application must still validate the operation and its result.
Connected ideas
Follow the engineering
Software, data and observability
HTTP status code
A protocol-level result code for an HTTP response. It helps classify outcomes, but application correctness and business success may require additional validation.
Connected ideas
Follow the engineering
Cloud, networking and security
HTTPS
HTTP protected by TLS in transit. Certificate validation and endpoint identity matter, but encrypted transport does not by itself authorize a request or prove the application is safe.
Connected ideas
Follow the engineering
Cloud, networking and security
Hyperscaler
A large cloud provider offering broad infrastructure and managed services. Service breadth can reduce operating work, but integration, commitments and data movement can also make a later migration expensive.
Connected ideas
Follow the engineering
Cloud, networking and security
Hypervisor
Software or a platform layer that manages virtual machines and their access to host resources. Device assignment, scheduling and isolation depend on the full supported setup.
Connected ideas
Follow the engineering
Cloud, networking and security
IAM
The policies and mechanisms governing which identities may access which resources. A role or credential should match a specific workload need instead of quietly broadening a migration’s authority.
Connected ideas
Follow the engineering
AI inference and model serving
ICI
A TPU interconnect used for communication between supported chips in a configured system. Its topology and communication cost matter when a workload spans multiple devices.
Connected ideas
Follow the engineering
Software, data and observability
Idempotency
A property that allows an operation to be repeated under its contract without producing an unintended additional effect. It is central to safe retries when the result of a previous attempt is uncertain.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Idempotency key
An identifier used to recognize repeated attempts at the same operation under a service contract. Its scope and retention rules must be defined so retries do not accidentally create a second effect.
Connected ideas
Follow the engineering
Business value and cost
Idle capacity
Paid or owned capacity that is not doing the useful work counted in the model. Some spare capacity is an intentional reliability choice; cost comparisons must distinguish that choice from avoidable waste.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Immutable
Not changed in place under the relevant contract. An immutable artifact gives references a stable meaning; new content should receive a new identity instead of silently replacing what a version or digest used to describe.
Connected ideas
Follow the engineering
AI inference and model serving
Inference
Running a trained model to produce outputs for new inputs. An inference service also has to handle scheduling, memory, failures, quality and latency around the model computation.
Connected ideas
Follow the engineering
Delivery and engineering tools
Infrastructure as code
Representing infrastructure configuration in versioned, reviewable source. The safety of a change still depends on its plan, credentials, ownership boundaries and verification.
Connected ideas
Follow the engineering
Kernels and numerical computing
INT4
A four-bit integer representation often used within quantized model storage or operations. Integer and floating-point four-bit schemes have different scaling and execution contracts.
Connected ideas
Follow the engineering
Hardware and execution
IOMMU
Hardware that translates and constrains device access to memory. It is an important part of device-assignment isolation, but it does not by itself define the complete security boundary of a virtual machine.
Connected ideas
Follow the engineering
Cloud, networking and security
IP
A network addressing and packet-delivery layer. An IP address identifies a network endpoint under its current routing context, not a durable application identity or permission.
Connected ideas
Follow the engineering
Delivery and engineering tools
JavaScript
A programming language used in browsers and many server environments. Interactive behavior should preserve accessible content and native navigation rather than making every explanation depend on script execution.
Connected ideas
Follow the engineering
Kernels and numerical computing
JAX
A numerical computing system that supports program transformations and accelerator execution. JAX tracing and compilation change where work happens; ordinary Python control flow is not automatically a GPU operation.
Connected ideas
Follow the engineering
Kernels and numerical computing
JIT
Compilation performed during program use rather than entirely ahead of deployment. Recompilation and warmup can affect latency, so separate compilation cost from steady-state execution deliberately.
Connected ideas
Follow the engineering
Software, data and observability
JSON
A structured data format used to exchange values such as objects and arrays. Valid JSON syntax does not establish that a payload obeys the application’s schema or is safe to act on.
Connected ideas
Follow the engineering
Kubernetes and production operations
Karpenter
A Kubernetes node-provisioning and lifecycle system. Its decisions depend on workload constraints and configured policy; it is not a substitute for correct requests, disruption planning or cost accounting.
Connected ideas
Follow the engineering
Kernels and numerical computing
Kernel fusion
Combining operations so that intermediate values can avoid extra launches or memory traffic. Fusion is useful only if the numerical contract remains correct and new register or scheduling costs do not outweigh the saving.
Connected ideas
Follow the engineering
Kubernetes and production operations
Kubelet
The node agent that coordinates the local execution of assigned Kubernetes workloads. Container runtime and device behavior remain separate layers below the scheduler’s placement decision.
Connected ideas
Follow the engineering
Kubernetes and production operations
Kubernetes
A system for managing declared application workloads across machines. It reconciles desired state, but still relies on correct application behavior, resource policies and operational decisions to deliver a useful service.
Connected ideas
Follow the engineering
Kubernetes and production operations
Kubernetes node
A machine registered with a Kubernetes cluster to run supported workloads. Scheduling uses declared resources and policy; a node is not automatically usable for a model until its devices and software are ready.
Connected ideas
Follow the engineering
AI inference and model serving
KV cache
Stored attention keys and values reused during language-model generation. Its size depends on the model’s attention layout, sequence lengths, data format and concurrency; weight size alone does not establish model fit.
Connected ideas
Follow the engineering
Cloud, networking and security
KVM
Linux virtualization infrastructure used with a supporting userspace stack to run virtual machines. It is one part of the guest, device and host isolation boundary, not a complete GPU hosting service by itself.
Connected ideas
Follow the engineering
Hardware and execution
L1 cache
A small, low-level cache close to execution resources. Its sharing rules and relationship with shared memory are architecture-specific, so an optimization must use the target GPU documentation rather than a generic diagram.
Connected ideas
Follow the engineering
Hardware and execution
L2 cache
A larger cache level that can serve work across a broader part of a processor. Reuse that fits and remains resident in L2 can reduce traffic to device memory, but a cache hit is not new DRAM bandwidth.
Connected ideas
Follow the engineering
Hardware and execution
Latency
The time taken by one operation or request. Average latency can hide slow requests, so a service usually needs explicit percentile targets and a clear boundary for what the timer includes.
Connected ideas
Follow the engineering
Software, data and observability
Lease
Time-bounded ownership or access under a system’s rules. Expiration, renewal and concurrent owners need explicit handling; a lease is not permanent authority.
Connected ideas
Follow the engineering
Cloud, networking and security
Least privilege
Giving an identity only the authority needed for its defined work. Time, resource and action scope matter, especially when an automated agent can turn a mistaken instruction into a real external effect.
Connected ideas
Follow the engineering
Cloud, networking and security
Linux
An operating-system kernel and its surrounding platform ecosystem. Linux manages process, memory and device boundaries; CUDA libraries and Kubernetes controllers act at different layers above it.
Connected ideas
Follow the engineering
Cloud, networking and security
Linux namespace
An operating-system mechanism that gives processes selected isolated views, such as process IDs or mounts. Namespaces do not replace resource limits, permissions or a complete sandbox policy.
Connected ideas
Follow the engineering
Kubernetes and production operations
Liveness
A check used to decide whether a running workload needs recovery such as restart. A dependency outage should not automatically turn a poorly chosen liveness check into a restart storm.
Connected ideas
Follow the engineering
AI inference and model serving
LLM
A model trained to process and generate language-like token sequences. Useful deployment depends on task quality, context and output lengths, memory use and service objectives—not only model size.
Connected ideas
Follow the engineering
Delivery and engineering tools
Locking
Coordination that prevents conflicting ownership of a resource or operation under a defined protocol. A lock’s scope, release and failure behavior matter; it is not a substitute for understanding which system should own the change.
Connected ideas
Follow the engineering
Hardware and execution
Logic gate
A circuit that maps input bits to a Boolean result. Gates are building blocks for arithmetic and control, but a small teaching circuit is not a complete model of a modern GPU datapath.
Connected ideas
Follow the engineering
Software, data and observability
Logs
Recorded events or messages from a system. Logs are useful when structured and correlated, but should not become a dumping ground for credentials or confidential payloads.
Connected ideas
Follow the engineering
Software, data and observability
Loki
A log aggregation system used to search log streams. Its data is logs, not a substitute for a properly defined metric or a distributed trace.
Connected ideas
Follow the engineering
AI inference and model serving
Machine learning
Methods that fit model behavior from data rather than expressing every decision as explicit rules. Deployment must connect the learned artifact with the data, evaluation and serving conditions it depends on.
Connected ideas
Follow the engineering
AI inference and model serving
Machine-learning model
A learned artifact that maps inputs to outputs under a particular training and evaluation setup. A model file alone is not a complete release: preprocessing, dependencies, configuration and serving behavior also matter.
Connected ideas
Follow the engineering
Cloud, networking and security
Managed service
A service for which a provider operates a defined part of the stack. The contract determines the actual boundary; application correctness, data recovery and customer obligations may remain yours.
Connected ideas
Follow the engineering
Kernels and numerical computing
Matrix multiplication
An operation combining rows and columns to produce a new matrix. Tiling changes how often operands move, not the mathematical product or the need to handle boundary shapes correctly.
Connected ideas
Follow the engineering
Model platforms and agent workflows
MCP
A protocol for connecting applications with tools and contextual resources. A standardized connection does not automatically authorize every exposed action or make returned content trustworthy.
Connected ideas
Follow the engineering
Hardware and execution
Memory
Storage that a running computation can access for instructions and data. Registers, caches, host RAM and GPU memory have different capacities, sharing rules and access costs; they are not one interchangeable pool.
Connected ideas
Follow the engineering
Hardware and execution
Memory controller
Hardware that coordinates requests to memory channels. Access patterns and competing requests influence the service it provides; a memory controller does not remove the cost of moving bytes.
Connected ideas
Follow the engineering
Kernels and numerical computing
Memory-bound
A workload currently constrained more by moving data than by performing arithmetic, under the chosen measurement boundary. Reducing traffic may help more than adding compute, but the diagnosis needs evidence.
Connected ideas
Follow the engineering
Software, data and observability
Message queue
A mechanism that buffers work or messages between producers and consumers. Delivery semantics, ordering, backlog and replay behavior matter as much as peak throughput.
Connected ideas
Follow the engineering
Software, data and observability
Metrics
Measurements aggregated over defined dimensions and time windows. A metric needs a clear unit and meaning; an average or count can conceal failures if its denominator or exclusions are unclear.
Connected ideas
Follow the engineering
Software, data and observability
Microservice
A service organized around a bounded responsibility and an explicit interface. Splitting code into more processes is not enough; data ownership, deployment and recovery boundaries must become clearer.
Connected ideas
Follow the engineering
Cloud, networking and security
MIG
A supported NVIDIA mechanism for partitioning some GPU resources into instances. Partition shape, isolation guarantees and device support must be checked; a slice is not an arbitrary fraction of every GPU resource.
Connected ideas
Follow the engineering
Model platforms and agent workflows
MLflow
Tools for organizing machine-learning experiments, artifacts and model lifecycle work. Its metadata and stored artifacts have different failure and recovery boundaries and should remain tied to an identifiable release.
Connected ideas
Follow the engineering
Model platforms and agent workflows
MLOps
Engineering practices for making machine-learning work reproducible, deployable and operable. The platform must connect data, models, evaluation, releases and recovery rather than stopping at a successful notebook run.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Model artifact
An identifiable output used by a build, experiment or release, such as weights or a compiled package. An artifact should be tied to its inputs and intended use instead of being mutable anonymous workstation state.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Model monitoring
Observing whether model inputs and useful behavior remain within the intended operating contract. Changes in data or outcomes require investigation; a healthy process metric alone does not establish model quality.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Model registry
A system for organizing model versions and lifecycle metadata. Registry entries still need to point to the correct artifacts and deployment contracts; a named version is not a complete production release.
Connected ideas
Follow the engineering
AI inference and model serving
Model training
The process of fitting model parameters from data. Training artifacts, input versions and evaluation evidence need to be identifiable so that a model can be reproduced and deployed safely.
Connected ideas
Follow the engineering
AI inference and model serving
Model weights
Learned numerical parameters used by a model. Weights are only part of runtime memory use; caches, intermediate activations, workspaces and captured execution state can also occupy memory.
Connected ideas
Follow the engineering
Software, data and observability
Monolith
An application deployed or operated as a larger combined unit. A monolith is not inherently wrong; splitting it should address a real ownership, delivery or scaling constraint.
Connected ideas
Follow the engineering
Kernels and numerical computing
MXFP4
A microscaling FP4 format family that groups low-precision values with scale information. Compare its actual representation and supported kernels rather than treating all four-bit methods as equivalent.
Connected ideas
Follow the engineering
Kubernetes and production operations
Namespace
A named isolation or organization boundary whose meaning depends on the system. Kubernetes namespaces organize API resources; Linux namespaces isolate selected operating-system views, and neither alone provides complete workload security.
Connected ideas
Follow the engineering
Cloud, networking and security
NAT
Translation between network address spaces along a traffic path. NAT can affect routing, visibility and cost; it is not equivalent to an application-level authorization boundary.
Connected ideas
Follow the engineering
AI inference and model serving
NCCL
NVIDIA’s library for collective communication across supported GPUs. Communication performance depends on topology and transport as well as the operation and data size.
Connected ideas
Follow the engineering
Cloud, networking and security
Neocloud
A newer or more specialized cloud provider, often focused on accelerator capacity. Evaluate support, topology, availability, security and recovery alongside the GPU-hour quote.
Connected ideas
Follow the engineering
Cloud, networking and security
Network policy
Rules describing permitted traffic at a defined network boundary. Enforcement depends on the platform and paths involved; network restrictions do not replace application authorization.
Connected ideas
Follow the engineering
Hardware and execution
NIC
The hardware interface connecting a machine to a network. Its usable throughput depends on the link, topology, protocol and application, not just the interface’s advertised rate.
Connected ideas
Follow the engineering
Kubernetes and production operations
Node
An individual machine or participant in a distributed system. In a diagram, a node may instead represent a logical component, so check whether the surrounding explanation describes physical capacity or a responsibility boundary.
Connected ideas
Follow the engineering
Kubernetes and production operations
Node drain
Moving eligible workloads off a node for maintenance or lifecycle changes. Draining must respect the relevant disruption and shutdown contracts; it is not identical to lowering a workload’s replica count.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Notebook
An interactive document that combines executable code with results and explanation. A successful notebook run is not a production release unless its data, dependencies, evaluation and serving behavior are made reproducible.
Connected ideas
Follow the engineering
Kernels and numerical computing
Nsight Compute
An NVIDIA profiler for detailed kernel analysis. Its metrics can reveal memory behavior, occupancy and execution limits, but profiling overhead is not part of a normal device-timing result.
Connected ideas
Follow the engineering
Kernels and numerical computing
Nsight Systems
An NVIDIA timeline-oriented profiling tool for application and system behavior. It helps connect CPU activity, transfers and GPU work before deciding that an individual kernel is the problem.
Connected ideas
Follow the engineering
Hardware and execution
NUMA
A system topology in which memory access cost depends on which processor and memory region are involved. CPU placement, device attachment and memory locality can affect a GPU application’s data path.
Connected ideas
Follow the engineering
Kernels and numerical computing
Numerical stability
How well an algorithm avoids amplifying rounding and range problems while producing its intended result. A faster implementation is not equivalent if it silently changes the required numerical behavior.
Connected ideas
Follow the engineering
Kernels and numerical computing
Numerical tolerance
An explicit bound on acceptable numerical differences in a test. Absolute and relative error capture different situations; choose tolerances from the problem rather than to make a failing result pass.
Connected ideas
Follow the engineering
Kernels and numerical computing
NVFP4
An NVIDIA low-precision representation and supported execution approach using FP4 data with associated scaling. Its storage and numerical contract differs from simply storing arbitrary values in four bits.
Connected ideas
Follow the engineering
Hardware and execution
NVIDIA
A vendor of GPUs, accelerator systems and supporting software. Product family names are not interchangeable specifications; compare the exact device, memory, interconnect and supported software configuration.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
NVIDIA Container Toolkit
Integration that makes supported NVIDIA device and library access available within container execution. It is one part of the driver, runtime and orchestration stack, not a model-capacity or security guarantee.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
NVIDIA driver
Software connecting supported application interfaces with NVIDIA devices and operating-system facilities. Driver lifecycle and compatibility are deployment responsibilities, not something a Kubernetes scheduler solves by allocating a GPU.
Connected ideas
Follow the engineering
Hardware and execution
NVL72
A rack-scale configuration designation involving 72 interconnected GPUs in a supported platform. Aggregate resources describe the system, not a single GPU allocation available to every process.
Connected ideas
Follow the engineering
Hardware and execution
NVLink
An interconnect for communication between supported NVIDIA devices. Software must still partition data and coordinate transfers; connected GPU memories do not become one automatic shared allocation.
Connected ideas
Follow the engineering
Hardware and execution
NVMe
A storage protocol designed for low-latency access to suitable non-volatile devices. Local NVMe can be useful for caches and staging, but local instance storage is not automatically durable or replicated.
Connected ideas
Follow the engineering
Hardware and execution
NVSwitch
Switching technology used to connect supported GPU peers through NVLink fabrics. Fabric bandwidth and topology are different constraints from each GPU’s memory bandwidth.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Object storage
Storage addressed as objects under a service’s API and durability contract. Object versioning, permissions and recovery need to be designed explicitly when it holds model artifacts or application data.
Connected ideas
Follow the engineering
Software, data and observability
Observability
The ability to investigate a system’s behavior using meaningful signals and context. Collecting large amounts of telemetry is not enough if operators cannot connect symptoms to useful decisions.
Connected ideas
Follow the engineering
Kernels and numerical computing
Occupancy
The fraction of supported execution residency occupied by active work, under an architecture’s definition. Registers, shared memory and block limits influence occupancy; maximum occupancy is not automatically maximum performance.
Connected ideas
Follow the engineering
Delivery and engineering tools
OCI
Standards for container-related formats and runtimes. A compatible image format does not by itself establish provenance, safe permissions or correct workload behavior.
Connected ideas
Follow the engineering
Cloud, networking and security
OIDC
An identity protocol layered on OAuth 2.0, commonly used to establish authenticated identity claims. Those claims still need to be validated and mapped to narrowly scoped authorization.
Connected ideas
Follow the engineering
Cloud, networking and security
On-premises
Infrastructure operated in a location under the organization’s control or responsibility. Ownership changes the cost and operating boundary; hardware purchase price does not include every staffing, power, recovery or refresh obligation.
Connected ideas
Follow the engineering
AI inference and model serving
OOM
A failure caused by an allocation or memory limit that the workload cannot satisfy. The limit may belong to a process, container, GPU or another boundary; free memory elsewhere does not necessarily resolve it.
Connected ideas
Follow the engineering
Software, data and observability
OpenTelemetry
A set of standards and tools for producing and moving telemetry such as traces, metrics and logs. Instrumentation, sampling and data handling still need to match the service’s operational and privacy requirements.
Connected ideas
Follow the engineering
Software, data and observability
OpenTelemetry Collector
A component that receives, processes and exports telemetry through configured pipelines. Capacity, filtering and failure behavior need to be designed so that the observability path does not become an outage amplifier.
Connected ideas
Follow the engineering
Business value and cost
OpEx
Spending associated with ongoing operation under applicable accounting rules. A low resource rate can still have high operating cost once staffing, support and resilience responsibilities are included.
Connected ideas
Follow the engineering
Business value and cost
Opportunity cost
The value of the best alternative use of a constrained resource. It can inform a decision, but it should not be disguised as an observed cash expense or guaranteed additional revenue.
Connected ideas
Follow the engineering
Hardware and execution
OR
A Boolean operation that is true when at least one input is true. A full-adder example uses OR to combine carry conditions; it does not imply that a GPU uses that exact physical circuit.
Connected ideas
Follow the engineering
Software, data and observability
OTLP
A protocol used to exchange OpenTelemetry data. Protocol compatibility does not decide which data is appropriate to collect or how long it should be retained.
Connected ideas
Follow the engineering
Software, data and observability
Outbox
A pattern that records an intended message alongside a related local data change. Delivery can still be repeated, so consumers and reconciliation need an explicit duplicate-handling contract.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
P6-B200
An AWS EC2 offering designation for a particular B200-based instance configuration. Read its own product table for host, GPU, network and storage boundaries rather than substituting figures from a DGX system.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
P6-B300
An AWS EC2 offering designation for a particular B300-based instance configuration. Provider-reported aggregate resources and per-device allocatable memory are different quantities.
Connected ideas
Follow the engineering
AI inference and model serving
Paged attention
A way of managing attention-cache storage in smaller blocks rather than requiring one large contiguous allocation for each sequence. It can reduce allocation waste, but does not eliminate the bytes required by active tokens.
Connected ideas
Follow the engineering
Hardware and execution
PCIe
A high-speed interconnect used to attach devices to a system. Link generation, lane count, topology and competing traffic influence transfers; device compute capability alone does not describe that path.
Connected ideas
Follow the engineering
Kubernetes and production operations
PDB
A Kubernetes policy that limits certain voluntary evictions under its availability rules. It is not a universal availability guarantee and does not directly veto HPA or Deployment replica-count reductions.
Connected ideas
Follow the engineering
Software, data and observability
Percentile
A value describing a position in a distribution, such as the latency below which 95 percent of observed requests fall. Percentiles depend on the population and window and should not be averaged as if they were ordinary samples.
Connected ideas
Follow the engineering
Kubernetes and production operations
Persistent volume
A Kubernetes abstraction for storage with a lifecycle that can differ from an individual Pod. Persistence, replication, backups and recovery still depend on the underlying system and application contract.
Connected ideas
Follow the engineering
Cloud, networking and security
PID
An identifier for a running process within a process namespace. PID values and visibility can differ across namespaces, so a number alone does not describe the isolation or ownership boundary.
Connected ideas
Follow the engineering
Kubernetes and production operations
Pod
Kubernetes’ scheduling unit for one or more closely related containers. A Pod being scheduled or running does not by itself prove that the application is ready to serve requests.
Connected ideas
Follow the engineering
Kernels and numerical computing
Precision
The representational detail and range available in a numerical format. Lower precision can reduce storage or improve supported compute throughput, but accuracy, scaling and accumulation must be evaluated together.
Connected ideas
Follow the engineering
AI inference and model serving
Prefill
The inference phase that processes the input sequence before generating subsequent output tokens. Its compute and memory behavior differs from token-by-token decode, so measure the two phases separately.
Connected ideas
Follow the engineering
AI inference and model serving
Prefix caching
Reuse of compatible previously computed input-prefix state. A benchmark must say whether prefixes are shared and warm, since cache reuse can change the work being measured.
Connected ideas
Follow the engineering
Kernels and numerical computing
Profiling
Measuring where a program spends work, time or resources. A profiler can perturb execution, so use it to explain a bottleneck and retain separate controlled timing for the result.
Connected ideas
Follow the engineering
Kernels and numerical computing
Program tracing
Capturing a program’s operations into a representation that a compiler can transform. In JAX, values, shapes and Python behavior influence tracing; it differs from tracing requests across services.
Connected ideas
Follow the engineering
Software, data and observability
Prometheus
A metrics system commonly used for time-series queries and alerting. Metric design, scrape behavior and label cardinality determine the quality and cost of the resulting operational signals.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Prompt injection
An attempt to make a model or agent treat untrusted content as authoritative instructions. The defense needs enforced tool and data boundaries, not only a request that the model ignore bad text.
Connected ideas
Follow the engineering
Kernels and numerical computing
PTX
An NVIDIA intermediate instruction representation used in the CUDA compilation path. It is not always the final machine code executed by the selected GPU.
Connected ideas
Follow the engineering
Business value and cost
PUE
A facility metric relating total data-center energy use to IT equipment energy use. PUE is not a GPU’s power draw or an application-efficiency score, and the measurement boundary matters.
Connected ideas
Follow the engineering
Delivery and engineering tools
Python
A general-purpose programming language widely used for automation and machine learning. Python source is not automatically accelerator execution; libraries and compilation layers determine where the expensive work runs.
Connected ideas
Follow the engineering
AI inference and model serving
Quantization
Representing values with a more constrained numerical encoding, often with scale information. Smaller storage or faster supported kernels are useful only if the required model quality and numerical behavior survive.
Connected ideas
Follow the engineering
AI inference and model serving
Queueing
Waiting for a resource or service slot before work can proceed. Queueing can dominate request latency even when the execution stage itself is fast.
Connected ideas
Follow the engineering
Model platforms and agent workflows
RAG
A workflow that retrieves relevant material and supplies it as context for a model’s response. Retrieval can improve grounding, but access control, freshness and evaluation still govern whether the answer is useful and safe.
Connected ideas
Follow the engineering
Cloud, networking and security
RAID
Storage arrangements that combine devices for selected performance or failure-handling properties. Usable capacity differs from raw drive capacity, and RAID is not a substitute for a backup or recovery plan.
Connected ideas
Follow the engineering
Hardware and execution
RAM
Working memory used by the CPU and operating system. Host RAM is separate from GPU memory, so an application may pay for transfers, mappings or staging even when both capacities look sufficient.
Connected ideas
Follow the engineering
Cloud, networking and security
RBAC
Authorization through roles and their permitted operations on resources. A role name is not evidence of least privilege; inspect the actual scope and permissions.
Connected ideas
Follow the engineering
Cloud, networking and security
RDMA
Supported mechanisms for transferring data between machines with reduced involvement in normal CPU copy paths. Security, registration, transport support and topology still constrain the actual result.
Connected ideas
Follow the engineering
Kubernetes and production operations
Readiness
Whether an instance satisfies the declared condition for receiving useful traffic. Readiness should reflect the service contract; a running process or successful scheduling decision is not sufficient on its own.
Connected ideas
Follow the engineering
Kubernetes and production operations
Reconciliation
The process of bringing observed state toward a declared target. It is not an instantaneous transaction; errors, ownership conflicts and convergence need to remain observable.
Connected ideas
Follow the engineering
Software, data and observability
Recovery
Restoring a service and its required data to a supported operating state after failure. Recovery needs tested procedures and explicit loss/time objectives, not only a nominally redundant deployment.
Connected ideas
Follow the engineering
Hardware and execution
Register
Very small storage used by execution units for live values. Register demand affects how many work items can reside on a processor, and spills can add memory traffic rather than providing free extra space.
Connected ideas
Follow the engineering
Kernels and numerical computing
Register spill
A situation where a value cannot remain in the intended register allocation and uses another storage path. Spills can add memory traffic, so an apparent arithmetic optimization may create a different bottleneck.
Connected ideas
Follow the engineering
Kubernetes and production operations
Replica
An instance of a workload or replicated data under a particular controller’s contract. A desired replica count is not the same as the number of healthy instances currently serving traffic.
Connected ideas
Follow the engineering
Kubernetes and production operations
ReplicaSet
A Kubernetes controller that maintains a requested set of Pod replicas. A Deployment normally manages ReplicaSets as part of its rollout behavior.
Connected ideas
Follow the engineering
Software, data and observability
Replication
Maintaining additional copies of data or state under a defined consistency model. Copies can improve resilience, but correlated failures and logical corruption still require recovery planning.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Reproducibility
The ability to repeat or reconstruct a result under a sufficiently specified set of conditions. Code alone may not be enough: inputs, dependencies, configuration, hardware and measurement boundaries can all matter.
Connected ideas
Follow the engineering
AI inference and model serving
Request rate
The rate at which requests arrive or complete at a defined boundary. Daily active users do not determine this rate without a workload model, and arrival rate can exceed a system’s sustainable completion rate.
Connected ideas
Follow the engineering
Kubernetes and production operations
Resource limit
A configured upper bound enforced at a resource boundary. CPU throttling and memory termination have different effects, so a single utilization chart does not explain all limit-related failures.
Connected ideas
Follow the engineering
Kubernetes and production operations
Resource request
Declared resource demand used by Kubernetes scheduling and related policy. Requests are not the same as current usage; understated requests can make placement look feasible without making the service reliable.
Connected ideas
Follow the engineering
Software, data and observability
REST
An architectural style commonly used for resource-oriented web APIs. An endpoint’s practical safety still depends on its documented semantics, authentication, authorization and error handling.
Connected ideas
Follow the engineering
Software, data and observability
Retry
Another attempt after an operation fails or its result becomes uncertain. Retries need bounds and safe effect semantics; unbounded retries can duplicate actions or turn overload into a larger outage.
Connected ideas
Follow the engineering
Business value and cost
Rightsizing
Adjusting resource allocations to the workload and its reliability requirements. Observed usage, declared requests and enforceable limits are different inputs, so shrinking a number on a dashboard is not enough.
Connected ideas
Follow the engineering
Business value and cost
ROI
A comparison between an investment and specified returns over a defined period. Different modeled benefits cannot simply be added together when they overlap or represent incompatible quantities.
Connected ideas
Follow the engineering
Kubernetes and production operations
Rollback
Returning to a previously supported application or configuration state after a change. Data migrations and external effects may not be reversible just because an old container image is still available.
Connected ideas
Follow the engineering
Kubernetes and production operations
Rollout
The process of replacing or introducing a deployed version. A safe rollout connects availability, readiness, observation and rollback rather than treating a controller’s progress status as sufficient evidence.
Connected ideas
Follow the engineering
Kernels and numerical computing
Roofline
A model comparing a workload’s compute demand with its memory traffic to estimate a limiting physical bound. It helps formulate a hypothesis, not certify a measured latency or speedup.
Connected ideas
Follow the engineering
Software, data and observability
RPC
A call to an operation running across a process or network boundary. A timeout can leave the result uncertain, so retries must account for whether the operation may already have taken effect.
Connected ideas
Follow the engineering
Software, data and observability
RPO
The maximum acceptable data-loss window under a recovery plan. It should be evaluated against actual replication and backup behavior rather than inferred from the presence of a backup job.
Connected ideas
Follow the engineering
Software, data and observability
RTO
The target time to restore the required service after a defined failure. Provisioning, restoring data, checking correctness and reconnecting dependencies can all contribute.
Connected ideas
Follow the engineering
Delivery and engineering tools
Rust
A systems language with mechanisms for enforcing many memory and ownership constraints. Language guarantees are valuable, but they do not automatically establish the safety of external services, unsafe interfaces or an entire deployment.
Connected ideas
Follow the engineering
Software, data and observability
Saga
A pattern for coordinating a multi-step workflow through local actions and compensations. A compensation is a business action with its own failure modes, not an automatic reversal of every external effect.
Connected ideas
Follow the engineering
Software, data and observability
Sampling
Selecting which observations to retain or inspect. Sampling can control telemetry cost, but it changes which incidents or requests remain visible and must be understood when interpreting evidence.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Sandbox
A bounded execution environment intended to constrain selected actions and resources. Filesystem, process, network, credentials and approval controls have separate enforcement responsibilities.
Connected ideas
Follow the engineering
Kernels and numerical computing
SASS
A common name for the device-specific machine instructions produced for NVIDIA GPUs. Inspecting final generated instructions can explain behavior that is not obvious from high-level source alone.
Connected ideas
Follow the engineering
Kubernetes and production operations
Scheduler
A component that assigns eligible work to available execution locations according to declared constraints. Kubernetes scheduling places Pods; it does not schedule individual CUDA warps or make a model’s memory fit correct.
Connected ideas
Follow the engineering
Software, data and observability
Schema
A declared structure and set of constraints for data or an interface. Schema compatibility must account for both producers and consumers during a rollout or migration.
Connected ideas
Follow the engineering
Software, data and observability
SDK
Libraries and tools intended to help applications use a platform or service. An SDK simplifies integration but does not remove the underlying service, security or versioning contract.
Connected ideas
Follow the engineering
Cloud, networking and security
seccomp
A Linux facility for restricting a process’s available system calls under a policy. It can reduce attack surface, but must be combined with the other execution and authority boundaries a workload requires.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
Service account
An identity used by software rather than an interactive human login. Its authority should follow the workload’s actual needs and remain separate from unrelated administrative access.
Connected ideas
Follow the engineering
Software, data and observability
Service boundary
The division of responsibility and ownership between services. A useful boundary makes interfaces, data changes, failures and releases easier to reason about rather than adding coordination without independence.
Connected ideas
Follow the engineering
Hardware and execution
Shared memory
On a CUDA GPU, storage explicitly shared by threads in a block. It can reduce repeated device-memory reads, but the program must manage ownership, synchronization and the limited capacity correctly.
Connected ideas
Follow the engineering
Cloud, networking and security
Shared responsibility
The division of security and operational duties between a provider and its customer. A managed platform does not transfer every duty, so identify who owns configuration, data, access and recovery.
Connected ideas
Follow the engineering
Kubernetes and production operations
SIGTERM
A conventional operating-system signal requesting process termination. Handling it correctly requires application-specific shutdown behavior; receiving the signal is not proof that useful work has drained.
Connected ideas
Follow the engineering
Hardware and execution
SIMT
An execution model in which many logical threads run through shared instruction scheduling. Branches and masks can change which lanes participate, so independent-looking thread code still has execution constraints.
Connected ideas
Follow the engineering
Delivery and engineering tools
SKU
A product or offering identifier used in a catalog. In cloud comparisons, an exact SKU or configuration helps prevent mixing prices and specifications from different services.
Connected ideas
Follow the engineering
Software, data and observability
SLA
A contractual service commitment with its own terms and remedies. An engineering SLO and a provider’s SLA are related ideas but are not interchangeable guarantees.
Connected ideas
Follow the engineering
Software, data and observability
SLI
A measurement of a specific service outcome, such as successful requests within a latency bound. Its definition and exclusions determine whether an SLO reflects the user experience.
Connected ideas
Follow the engineering
Software, data and observability
SLO
A defined target for a service outcome over a stated measurement window. An SLO needs a meaningful indicator and population; a target alone is not evidence that the service achieved it.
Connected ideas
Follow the engineering
Kernels and numerical computing
Softmax
A normalization that converts a row of scores into values summing to one. A stable implementation subtracts a suitable reference before exponentiation and must handle the stated numerical domain correctly.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Software agent
Software that takes actions for a defined task. In AI workflows a model may help choose steps, but that does not grant the authority to use every tool, access every resource or repeat an uncertain external action.
Connected ideas
Follow the engineering
Software, data and observability
Span
A timed unit of work in a distributed trace, with relationships and selected attributes. Span boundaries and propagated context determine whether the resulting trace is useful for diagnosing a real request.
Connected ideas
Follow the engineering
Software, data and observability
SQL
A language family for querying and changing relational data. Transactions, isolation and schema changes determine the operational guarantees of a particular database workflow.
Connected ideas
Follow the engineering
Software, data and observability
SRE
An engineering approach to operating services with explicit reliability goals and feedback. Reliability work connects user outcomes, risk, change policy and measured service behavior.
Connected ideas
Follow the engineering
Hardware and execution
SSD
Storage built from solid-state media rather than rotating disks. Latency, endurance, throughput and persistence requirements all matter; the label SSD does not imply every storage path performs alike.
Connected ideas
Follow the engineering
Cloud, networking and security
SSH
A protocol commonly used for authenticated remote access and command execution. Possession of an SSH path is still scoped authority, not permission to inspect unrelated data or perform every operation on the host.
Connected ideas
Follow the engineering
Cloud, networking and security
SSRF
A class of vulnerability where an attacker influences a server to make unintended network requests. Destination validation, routing boundaries and least privilege matter, particularly for tools that fetch user-supplied URLs.
Connected ideas
Follow the engineering
Business value and cost
Staff capacity
Time or effort made available for other work. Its value depends on whether the organization can use that capacity; it is not automatically a reduction in payroll or a cash saving.
Connected ideas
Follow the engineering
Kubernetes and production operations
Startup probe
A Kubernetes probe that allows a workload’s initialization phase to complete under its configured budget before other health behavior takes over. It must account for real startup work rather than hiding a permanent failure.
Connected ideas
Follow the engineering
Kubernetes and production operations
StatefulSet
A Kubernetes controller for workloads that need stable identity and related lifecycle guarantees. It does not automatically solve the application’s replication, backup or recovery problems.
Connected ideas
Follow the engineering
Hardware and execution
Streaming multiprocessor
A group of scheduling, execution and memory resources inside an NVIDIA GPU. The number of resident warps and blocks depends on resources such as registers and shared memory, not just the device headline.
Connected ideas
Follow the engineering
Cloud, networking and security
Subnet
A subdivision of an IP network with defined address and routing behavior. Subnet placement alone is not an authorization policy or a guarantee of workload isolation.
Connected ideas
Follow the engineering
Hardware and execution
SXM
A module form factor used by certain NVIDIA accelerators. Power, cooling and interconnect characteristics depend on the supported platform; SXM and PCIe product figures should not be mixed casually.
Connected ideas
Follow the engineering
Kernels and numerical computing
Synchronization
Coordination that establishes when work or memory effects are safe to observe. Synchronization may be required for correctness and measurement, but it can also introduce waiting or reduce overlap.
Connected ideas
Follow the engineering
Kubernetes and production operations
Taint
A Kubernetes node property used to repel Pods unless the corresponding placement rules permit them. Taints help express placement intent but are not a substitute for device isolation or authorization.
Connected ideas
Follow the engineering
Kernels and numerical computing
TB
A decimal terabyte is one trillion bytes; TB/s is a transfer rate. State whether a figure covers one device, an aggregate fabric or a different system boundary.
Connected ideas
Follow the engineering
Business value and cost
TCO
The cost of owning or using a system over a defined scope and period. Include idle capacity, operating work, commitments, recovery and exit costs rather than comparing only the quoted resource rate.
Connected ideas
Follow the engineering
Cloud, networking and security
TCP
A transport protocol providing an ordered byte stream with its own reliability mechanisms. Application timeouts, retries and semantic success remain separate concerns above that transport.
Connected ideas
Follow the engineering
Software, data and observability
Tempo
A distributed-trace backend. Trace records help investigate specific request paths; they should not be presented as if Tempo stores the service’s metric time series.
Connected ideas
Follow the engineering
Hardware and execution
Tensor Core
Specialized hardware for supported matrix operations and number formats. Peak Tensor Core throughput applies to specific conditions; arbitrary scalar code does not automatically achieve it.
Connected ideas
Follow the engineering
AI inference and model serving
Tensor parallelism
Splitting parts of a model’s tensor operations across devices. Communication and synchronization become part of latency and throughput, and each device still has its own allocation boundary.
Connected ideas
Follow the engineering
AI inference and model serving
TensorRT-LLM
NVIDIA software for optimized language-model inference on supported hardware and models. Engine version, build settings and numerical formats must be part of a reproducible comparison.
Connected ideas
Follow the engineering
Delivery and engineering tools
Terraform
Infrastructure-as-code tooling for planning and managing supported resources. State, locking, reviewed plans and ownership define its safety boundary; application rollout responsibilities may belong elsewhere.
Connected ideas
Follow the engineering
Delivery and engineering tools
Terraform state
Terraform’s recorded mapping between configuration and managed resources. Protecting and coordinating state access is different from controlling application data or a model’s runtime state.
Connected ideas
Follow the engineering
Delivery and engineering tools
Testing
Exercising a system against an observable contract and checking the result. Useful tests expose plausible failures and boundaries instead of merely asserting that internal fields were copied.
Connected ideas
Follow the engineering
Kernels and numerical computing
TF32
An NVIDIA Tensor Core compute format with particular precision and range behavior. It does not mean an arbitrary Float32 program receives identical arithmetic at a quoted Tensor Core peak.
Connected ideas
Follow the engineering
Kernels and numerical computing
TFLOP
A trillion floating-point operations, or a rate when written per second. Peak rates depend on precision, operation type and hardware conditions and are not automatically achieved application throughput.
Connected ideas
Follow the engineering
Kernels and numerical computing
Thread
A logical sequence of execution. CUDA threads have defined ownership and synchronization rules and are scheduled in warps; they are not independent miniature CPUs.
Connected ideas
Follow the engineering
Kernels and numerical computing
Thread block
A group of CUDA threads that can cooperate through supported block-level synchronization and shared memory. Blocks must not depend on an unspecified execution order across the grid.
Connected ideas
Follow the engineering
Kubernetes and production operations
Throttling
Restricting how quickly work can consume a resource. Throttling may protect a budget, but it can also increase latency even when a broader host-level utilization average appears low.
Connected ideas
Follow the engineering
Hardware and execution
Throughput
The amount of work completed per unit of time. Useful throughput must say which work succeeded and met its quality and latency requirements, rather than counting every attempted request.
Connected ideas
Follow the engineering
Kernels and numerical computing
Tiling
Dividing a computation into reusable blocks of data and work. Tiling can reduce repeated memory traffic, but it adds coordination and resource demands that must be tested on the target workload.
Connected ideas
Follow the engineering
Cloud, networking and security
TLS
A protocol for authenticating and protecting supported network connections. It protects a transport boundary; application permissions and data-handling rules still need to be enforced.
Connected ideas
Follow the engineering
Kernels and numerical computing
TMA
A Blackwell/Hopper-era NVIDIA facility for supported asynchronous tensor-memory transfers. Its use still requires the correct layout, synchronization and supported instruction path.
Connected ideas
Follow the engineering
AI inference and model serving
Token
A unit whose meaning depends on context. Language models process pieces of text called tokens, while authentication tokens carry access authority; their limits, prices and security properties are not interchangeable.
Connected ideas
Follow the engineering
AI inference and model serving
Tokenization
Converting text into the token sequence a language model uses. Tokenization rules affect input length and workload shape, and the CPU-side work can contribute to end-to-end latency.
Connected ideas
Follow the engineering
Kubernetes and production operations
Toleration
A Kubernetes Pod rule that allows it to tolerate a matching node taint. It permits consideration of that node; it does not force placement or reserve a particular GPU.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Tool calling
A model-driven workflow selecting a structured operation for another system to execute. Arguments, permissions, approval and result handling must be enforced outside the model’s suggestion.
Connected ideas
Follow the engineering
AI inference and model serving
TPOT
A measure of the time associated with producing output tokens after generation begins. Its definition and aggregation must be explicit; good TTFT does not guarantee smooth or fast decode.
Connected ideas
Follow the engineering
AI inference and model serving
TPU
Google’s accelerator family for supported tensor workloads. Comparing a TPU with a GPU requires the exact configuration, software path and workload boundary, not unrelated vendor peak figures.
Connected ideas
Follow the engineering
Software, data and observability
Transaction
A group of operations governed by specific commit and failure semantics. Local database transactions do not automatically make a multi-service workflow atomic.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Trust boundary
A point where data, instructions or authority cross between differently trusted components. A system should validate and constrain the crossing instead of letting a convincing format or model output silently raise trust.
Connected ideas
Follow the engineering
AI inference and model serving
TTFT
Time from the chosen request start boundary until the first output token is available. Queueing, model startup and prefill may contribute, so state which stages a TTFT measurement includes.
Connected ideas
Follow the engineering
Delivery and engineering tools
TypeScript
A language that adds a static type system to JavaScript development. Types help with some contracts, but runtime inputs and external systems still require validation.
Connected ideas
Follow the engineering
Cloud, networking and security
UDP
A transport protocol for datagrams without TCP’s stream guarantees. The application or another protocol must define any required ordering, reliability and loss handling.
Connected ideas
Follow the engineering
Cloud, networking and security
UID
A user identifier within an operating-system or platform identity domain. A numeric UID only has meaning within that domain and its mappings; it is not a universal identity or permission.
Connected ideas
Follow the engineering
Software, data and observability
URI
An identifier for a resource under a defined syntax. Some URIs identify network locations, while others name resources without implying that a browser should fetch them.
Connected ideas
Follow the engineering
Software, data and observability
URL
An address describing where and how to access a resource. Parse and validate URLs by their actual scheme, host and path rather than trusting their visual appearance or a substring.
Connected ideas
Follow the engineering
Business value and cost
USD
A currency designation for United States dollars. Cost examples still need a time period, unit and accounting boundary; a dollar figure alone is not a provider quote or a verified saving.
Connected ideas
Follow the engineering
Business value and cost
Utilization
A measurement of activity under a particular resource definition. High utilization is not automatically good service, and low utilization does not by itself prove that capacity can be removed safely.
Connected ideas
Follow the engineering
Software, data and observability
Validation
Checking an input or result against a stated contract. A successful parser or transport response is not enough if semantic limits, permissions or required invariants have not been checked.
Connected ideas
Follow the engineering
Hardware and execution
vCPU
A virtual processor presented to a virtual machine. A vCPU may share physical execution resources; its count is not a count of GPU cores or a portable performance guarantee across cloud instance types.
Connected ideas
Follow the engineering
Model platforms and agent workflows
Vector search
Finding items using similarity between numerical representations. Filtering, access controls and relevance evaluation still matter; the nearest vector is not automatically the right answer for the task.
Connected ideas
Follow the engineering
Delivery and engineering tools
Versioning
Identifying and controlling which revision of source, data or a dependency is in use. A version label is useful when it resolves to an unambiguous artifact and a compatible contract.
Connected ideas
Follow the engineering
Cloud, networking and security
VFIO
A Linux framework for controlled userspace access to devices, commonly used in supported passthrough setups. Correct device grouping and IOMMU configuration remain part of the security boundary.
Connected ideas
Follow the engineering
Cloud, networking and security
vGPU
A virtualized GPU offering under a particular vendor and platform contract. Sharing, scheduling and isolation differ by implementation, so it should not be equated automatically with passthrough or MIG.
Connected ideas
Follow the engineering
Cloud, networking and security
Virtual machine
A virtualized machine with an operating-system and resource boundary supplied by a virtualization stack. Its devices, CPU scheduling and memory remain subject to the host and platform configuration.
Connected ideas
Follow the engineering
AI inference and model serving
vLLM
An open-source language-model serving engine. Compare its exact version, model support and workload configuration against other engines under the same latency and quality contract.
Connected ideas
Follow the engineering
Kubernetes and production operations
VPA
Kubernetes automation for recommendations or changes to workload resource sizing, depending on mode and setup. Resizing policy, workload disruption and interactions with other autoscaling controls need explicit ownership.
Connected ideas
Follow the engineering
Cloud, networking and security
VPC
A logically configured private networking boundary in a cloud platform. Routing, access rules, DNS and service integrations still determine which systems can actually reach one another.
Connected ideas
Follow the engineering
Hardware and execution
VRAM
A common label for memory associated with a graphics processor. Read the exact device and memory technology behind the number: capacity, bandwidth and allocatable free space are different quantities.
Connected ideas
Follow the engineering
Platforms and implementation boundaries
W3C
A standards organization associated with web technologies. A shared standard can improve interoperability, but an implementation still needs to meet the application’s accessibility, correctness and security requirements.
Connected ideas
Follow the engineering
Kernels and numerical computing
Warmup
Preparation before measurement so that effects such as compilation and allocation are handled deliberately. Warmup is not a reason to hide cold-start behavior when users actually experience it.
Connected ideas
Follow the engineering
Hardware and execution
Warp
A group of 32 logical CUDA lanes whose instructions are scheduled together. Only some lanes may participate in a particular instruction; lane activity is not the same thing as whole-GPU utilization.
Connected ideas
Follow the engineering
Delivery and engineering tools
WebAssembly
A portable low-level execution format supported by modern browsers and other runtimes. Here small WASM modules calculate teaching models and interaction geometry; they do not run a remote GPU or compile CUDA in the browser.
Connected ideas
Follow the engineering
Cloud, networking and security
Workload identity
A way to give a running workload a verifiable identity and obtain scoped access without treating long-lived copied secrets as the default. Trust mappings and resource permissions remain explicit parts of the design.
Connected ideas
Follow the engineering
Kernels and numerical computing
XLA
A compiler infrastructure used to optimize and execute supported tensor computations. Compiler decisions, shapes and generated operations matter when comparing a high-level program with a custom kernel.
Connected ideas
Follow the engineering
Hardware and execution
XOR
A Boolean operation that is true when its two input bits differ. XOR helps form the sum bit in a full adder, while carry logic handles the additional place value.
Connected ideas
Follow the engineering
Software, data and observability
YAML
A data-serialization format often used for configuration. The parsed values and consuming API contract matter more than whether a file looks like a familiar deployment example.