agent 推論 · GPU 效率

UnieInfra —— agent 推論,優化到每個 token

UnieInfra 是為 AI agent 而非僅是 LLM 呼叫調校的推論平台 —— 低負載下吞吐量最高 4×,高並發下延遲降低 2×,並以自動參數調校在你既有的硬體上最大化吞吐量,支援 AMD、Nvidia、Qualcomm 與 Intel。

WITH UNIEINFRA
UnieInfra

UnieInfra

Scheduler · Kernels · Autotune

GPU Server 01+ UNIEINFRA
throughput11,487 tok/s
GPU Server 02+ UNIEINFRA
throughput11,226 tok/s
GPU Server 03+ UNIEINFRA
throughput12,103 tok/s

低負載下的吞吐量,加速 AI 回應

高並發下延遲降低(TTFT)

96%

生產部署中的 GPU 記憶體使用率

4

加速器生態:AMD · Nvidia · Qualcomm · Intel

同時兼顧高吞吐量與低 time-to-first-token。

Total throughput

higher is better

Qwen3.5-122B-A10B (FP16) · Nvidia H200 × 2 · test by InferenceMAX.

UnieInfravLLM 0.21.0
4.5×
4.7×
4.6×
3.25×
vLLM crash
1163264128256

concurrency

Throughput vs. time-to-first-token

up & right is better

UnieInfra holds high throughput while collapsing TTFT as concurrency rises.

UnieInfravLLM 0.21.0
2000160012008004000Time to first token (ms)Total throughput (token/sec)c=64c=32c=16c=1c=64c=32c=16c=1

為何推論決定了 agent 的經濟性。

更高吞吐量

低負載下吞吐量最高 2× —— AI 回應更快。

更低延遲

高並發下延遲降低 2× —— 等待更少。

自動參數調校

依你的硬體自動調校批次、排程與記憶體。

核心 (kernel) 優化

調校過的 kernel 榨出每張加速卡的效能。

多硬體部署

單一引擎橫跨 AMD、Nvidia、Qualcomm、Intel。

私有部署

在你的資料中心內運行,完全掌控。

一套推論引擎,跨四大加速器平台

NVIDIA
AMDQualcommIntel

讓 agent 變得可負擔的推論

當一個 agent 每個任務要做許多次模型呼叫時,吞吐密度就決定了成本與可擴展性。UnieInfra 正是為這種工作負載而打造。

  • 規模化下更低的 token 成本
  • 高並發下穩定
  • 在你既有的硬體上運行
4× open-source stack
=
+UnieInfra
1 rack · same throughput

Token-efficient throughput density means the work of four racks on a stock open-source stack runs on a single rack with UnieInfra.

用你的工作負載實測 UnieInfra。