| Application Pattern |
SingleStream |
Processes one query at a time and measures the time required to complete an individual inference request. |
Median latency, 90th-percentile latency, 99th-percentile latency, and completed queries. |
Useful for interactive applications where each user request requires a fast and predictable response. |
Required |
| Application Pattern |
MultiStream |
Processes multiple related samples together while maintaining a defined response-time limit for the stream. |
Queries per stream, stream latency, batch composition, and end-to-end response time. |
Relevant to synchronized sensor, video, or other multi-input workloads that need consistent timing. |
Workload-Dependent |
| Application Pattern |
Server |
Processes requests arriving at variable times and evaluates performance under a defined request-arrival pattern. |
Queries per second, 99th-percentile latency, request rate, queue depth, and rejected or delayed requests. |
Represents online services where throughput and tail-latency consistency must be balanced. |
High |
| Application Pattern |
Offline |
Processes a known dataset without interactive response-time requirements. |
Samples per second, total completion time, batch size, and energy consumed per sample. |
Suitable for batch analytics, media processing, indexing, and other jobs optimized for maximum throughput. |
High |
| Latency |
Average Latency |
The arithmetic mean time from request submission to inference completion. |
Milliseconds per query, measured separately for warm-up and steady-state execution. |
Provides a general performance view, but can hide occasional slow requests. |
Secondary |
| Latency |
Tail Latency |
A high-percentile latency measurement, commonly reported at the 90th, 95th, or 99th percentile. |
Milliseconds per query at each selected percentile, including software and data-transfer overhead. |
Critical for service-level objectives because a small number of slow requests can affect user experience. |
Required |
| Throughput |
Queries Per Second |
The number of completed inference queries per second during a defined measurement interval. |
Steady-state queries per second, concurrency level, batch size, and request-arrival rate. |
Shows whether a server can support expected traffic without excessive queuing or latency degradation. |
Required |
| Throughput |
Samples Per Second |
The number of individual input samples processed per second, including samples grouped into batches. |
Samples per second, effective batch size, total runtime, and dataset size. |
Important for offline processing and for comparing performance when one query contains multiple samples. |
High |
| Scalability |
Concurrency Scaling |
The change in latency and throughput as the number of simultaneous requests increases. |
Results at low, medium, and high concurrency; throughput curve; latency curve; queueing time. |
Reveals whether performance remains stable when traffic increases and helps identify saturation points. |
Required |
| Model Configuration |
Precision Mode |
The numerical representation used during inference, such as FP32, FP16, BF16, or INT8. |
Precision, calibration method where applicable, accuracy result, latency, and throughput. |
Lower-precision execution can improve efficiency, but it must preserve application-level accuracy. |
Required |
| Model Configuration |
Batch Size |
The number of samples processed together in one inference operation. |
Batch size, latency per batch, samples per second, memory use, and maximum sustainable batch size. |
Batching often increases throughput but may add waiting time and memory pressure. |
Required |
| Memory Capacity |
Peak Accelerator Memory |
The maximum memory consumed by model weights, activations, runtime buffers, and input data. |
Peak memory usage in GB, unused capacity, model size, and memory usage at target batch size. |
Determines whether the server can run the intended model and future model versions without memory-related failures. |
Required |
| Interconnect |
Host-to-Accelerator Transfer |
The time and bandwidth required to move input and output data between the host system and the inference accelerator. |
Transfer latency, effective bandwidth, data-transfer overlap, and percentage of total inference time. |
Data movement can become the bottleneck when models are small, inputs are large, or multiple accelerators are used. |
Workload-Dependent |
| Multi-Accelerator Scaling |
Parallel Efficiency |
The achieved multi-accelerator performance divided by the ideal linear performance increase. |
Single-accelerator throughput, multi-accelerator throughput, scaling factor, and communication overhead. |
Helps determine whether adding accelerators will deliver practical performance gains. |
High |
| Accuracy |
Application Accuracy |
The quality of model outputs under the selected precision, preprocessing, and optimization settings. |
Task-specific metric, reference result, optimized result, accuracy difference, and acceptance threshold. |
A high-throughput server is unsuitable if optimization causes unacceptable prediction-quality loss. |
Required |
| Energy Efficiency |
Energy Per Query or Sample |
The energy consumed to complete one inference query or process one input sample. |
Joules per query or sample, average power, peak power, throughput, and measurement duration. |
Supports operating-cost analysis and helps compare systems under equivalent workload and accuracy conditions. |
High |
| Reliability |
Sustained Performance |
The ability to maintain stable latency and throughput during an extended workload run. |
Performance over time, thermal state, throttling events, error rate, and service interruptions. |
Short benchmark bursts may not reveal thermal throttling, memory leaks, or long-run stability problems. |
Required |
| Total Cost of Ownership |
Performance per Watt and Performance per Cost |
Inference output normalized by energy use or system acquisition and operating cost. |
Queries per second per watt, samples per second per currency unit, maintenance cost, and expected service life. |
Provides a more useful purchasing comparison than peak performance alone. |
High |