NVIDIA announced that TensorRT Edge-LLM is the first solution to complete the MLPerf Edge Agentic Benchmark, running up to 6.4x faster on Jetson AGX Thor, making on-device AI agents practical for edge hardware.

1 min read

TensorRT Edge-LLM Tops MLPerf Edge Agentic Benchmark at 6.4x Faster on Jetson AGX Thor

What NVIDIA announced

FAQ

What is TensorRT Edge-LLM?

It is NVIDIA's framework for optimizing large language model inference on edge devices like Jetson AGX Thor, designed for multi-step agent workflows rather than single-prompt responses.

What is the MLPerf Edge Agentic Benchmark?

A new benchmark from MLCommons that simulates an AI agent running on edge hardware: selecting tools, evaluating results, and continuing reasoning in a long context, instead of answering a single query.

How does this benefit MENA enterprises?

It enables running AI agents locally in industrial robots, vehicles, and remote sites without relying on cloud latency or stable connectivity, while improving privacy and compliance.

Does this replace the cloud?

No — it enables a hybrid model where latency-sensitive tasks run on-device while heavy workloads or training remain in the cloud.

Source: NVIDIA Developer (AI)

AI-assisted content, human-reviewed.