Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
NVIDIA releases TensorRT Model Connect in public preview for streamlined AI inference
00

NVIDIA releases TensorRT Model Connect in public preview for streamlined AI inference

Aug 18, 2026

NVIDIA has released TensorRT Model Connect in public preview, allowing developers to convert Hugging Face checkpoints to native C++ TensorRT inference in two commands without ONNX. The tool produces a versioned `.bundle` artifact that runs without PyTorch, targeting edge applications like robotics and medical devices. Built using OpenAI Codex agents, the project currently offers Linux aarch64 wheels, while x86_64 users must build from source.

TensorRT Model Connect release

  • ▪TensorRT Model Connect converts supported Hugging Face or local checkpoints to end-to-end TensorRT inference in two commands without an intermediate ONNX export step.
  • ▪NVIDIA released TensorRT Model Connect in public preview as an open-source project under the Apache-2.0 license on or before August 18, 2026.

Bundle artifact architecture

  • ▪The TensorRT Model Connect `.bundle` artifact splits build and runtime, using Python for engine construction while native profiles execute inference in C++.
  • ▪TensorRT Model Connect produces a versioned `.bundle` artifact running through native C++ task APIs, enabling inference execution without PyTorch in the runtime path.

Linux aarch64 platform requirements

  • ▪Users of TensorRT Model Connect on x86_64 platforms must utilize a Docker source-build path because x86_64 wheels are not published.
  • ▪Release wheels for TensorRT Model Connect target Linux aarch64 only, requiring Python 3.10 or 3.12, glibc 2.39 or newer, and TensorRT 11.1.0.106.

Robotics edge deployment applications

  • ▪TensorRT Model Connect targets teams with in-house inference stacks in industries like robotics, automotive in-vehicle compute, medical devices, and defense edge systems.
  • ▪TensorRT Model Connect supports C++ binary applications including on-device text generation, speech recognition, document parsing, embeddings, and diffusion image generation.

AI-assisted development process

  • ▪The July 29, 2026 GB300 snapshot of TensorRT Model Connect covers 105 profiles across 76 families, with 102 profiles exceeding their reference by over 5%.
  • ▪NVIDIA developed the entire TensorRT Model Connect project using OpenAI Codex agents under human direction and review.

1 source

Marktechpost
NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands
View source article

Featured stories

View more in AI for developers

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

TSMC weighs multibillion-dollar Texas campus expansion for AI chips

Oct 1, 2026 · 2 sources

Story comments

Loading comments…

Related Projects

Hugging FaceNvidia

Topics

AI for developersCompute, chips & AI infrastructureAI tools & productsAI inference (scaling)Open-source AI

Featured stories

View more in AI for developers

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

TSMC weighs multibillion-dollar Texas campus expansion for AI chips

Oct 1, 2026 · 2 sources