NVIDIA has released TensorRT Model Connect in public preview, allowing developers to convert Hugging Face checkpoints to native C++ TensorRT inference in two commands without ONNX. The tool produces a versioned `.bundle` artifact that runs without PyTorch, targeting edge applications like robotics and medical devices. Built using OpenAI Codex agents, the project currently offers Linux aarch64 wheels, while x86_64 users must build from source.
TensorRT Model Connect release
- ▪TensorRT Model Connect converts supported Hugging Face or local checkpoints to end-to-end TensorRT inference in two commands without an intermediate ONNX export step.
- ▪NVIDIA released TensorRT Model Connect in public preview as an open-source project under the Apache-2.0 license on or before August 18, 2026.
Bundle artifact architecture
- ▪The TensorRT Model Connect `.bundle` artifact splits build and runtime, using Python for engine construction while native profiles execute inference in C++.
- ▪TensorRT Model Connect produces a versioned `.bundle` artifact running through native C++ task APIs, enabling inference execution without PyTorch in the runtime path.
Linux aarch64 platform requirements
- ▪Users of TensorRT Model Connect on x86_64 platforms must utilize a Docker source-build path because x86_64 wheels are not published.
- ▪Release wheels for TensorRT Model Connect target Linux aarch64 only, requiring Python 3.10 or 3.12, glibc 2.39 or newer, and TensorRT 11.1.0.106.
Robotics edge deployment applications
- ▪TensorRT Model Connect targets teams with in-house inference stacks in industries like robotics, automotive in-vehicle compute, medical devices, and defense edge systems.
- ▪TensorRT Model Connect supports C++ binary applications including on-device text generation, speech recognition, document parsing, embeddings, and diffusion image generation.
AI-assisted development process
- ▪The July 29, 2026 GB300 snapshot of TensorRT Model Connect covers 105 profiles across 76 families, with 102 profiles exceeding their reference by over 5%.
- ▪NVIDIA developed the entire TensorRT Model Connect project using OpenAI Codex agents under human direction and review.
Story comments
Loading comments…