Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
H Company releases NeoMME multimodal encoders for visual document retrieval
00

H Company releases NeoMME multimodal encoders for visual document retrieval

Sep 6, 2026

H Company has released NeoMME, a family of 260M and 800M single-tower bidirectional multimodal encoders designed to optimize visual document retrieval. By eliminating separate vision towers and causal decoders, NeoMME processes text and raw image patches within a single Transformer. The 260M model achieves 0.523 nDCG@10 on ViDoRe v3, matching the performance of much larger models while offering high throughput and compression options.

NeoMME bidirectional encoder architecture

  • ▪The NeoMME tokenizer is a whitespace-unconstrained BPE with a 131,072-entry vocabulary that emits 44.4% fewer tokens than ModernBERT across 14 target languages in FLORES-200 devtest.
  • ▪H Company released NeoMME, a family of 260M and 800M parameter single-tower bidirectional multimodal encoders that process text and raw image patches through a single Transformer.
  • ▪Both NeoMME models support a 16,384-token context window and utilize symmetric sliding-window attention on most layers, with global attention on every sixth layer and the final layer.
  • ▪The NeoMME architecture processes text via an ALBERT-style factorized embedding and raw 32x32 RGB image patches via a 2-layer MLP, omitting a separate vision tower and causal decoder.

Masked diffusion pretraining approach

  • ▪A cross-modal ablation probe showed that at 90% masking, visible page patches raised masked-token accuracy by 38.4 points for the NeoMME-260M model and 40.5 points for the NeoMME-800M model.
  • ▪NeoMME models are pretrained using discrete masked diffusion over text, with multimodal segments drawing a corruption rate of 0.30 to 1 to force the model to read the page.
  • ▪Pretraining of the NeoMME models processed approximately 524 billion packed input tokens, including roughly 290 billion text-only tokens, on NVIDIA H100 accelerators.

Visual document retrieval performance

  • ▪On the ViDoRe v3 benchmark, the NeoMME-Retriever-260M model scored 0.523 nDCG@10, which is within 0.002 of the 3.75B-parameter ColQwen2.5-v0.2 model.
  • ▪On the ViDoRe v3 benchmark, the NeoMME-Retriever-800M model scored 0.556 nDCG@10, landing 0.9 points behind the similarly sized Vultron Retriever Flash.
  • ▪On ViDoRe v1 and v2 benchmarks, the NeoMME models achieved nDCG@5 scores of 0.860 and 0.522 for the 260M model, and 0.874 and 0.559 for the 800M model.

Text retrieval limitations

  • ▪The authors of NeoMME attributed its weaker text retrieval performance to supervision scale, as NeoMME was trained on 430,000 text query examples compared to mLateOn's 660 million.
  • ▪On the BEIR-15 text retrieval benchmark, NeoMME late-interaction scored 0.4881 for the 260M model and 0.5126 for the 800M model, compared to 0.5722 for the 149M-parameter LateOn model.

Index compression techniques

  • ▪Applying hierarchical token pooling at factor 10 with int8 queries and documents reduced the NeoMME index size to 39.0 kB per page while retaining 99.16% of baseline nDCG@10.
  • ▪Standard late-interaction indexing for NeoMME yields approximately 1.5 MB per ViDoRe v3 document page in float32 format.

Encoding throughput benchmarks

  • ▪The NeoMME-260M model indexes 51.3 pages per second on a single NVIDIA L40S accelerator at a matched 2048x2048 input, which is 1.97 times faster than ColModernVBERT.
  • ▪The NeoMME-260M model encodes a query in 78.3 milliseconds on a CPU-only host.

1 source

Marktechpost
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
View source article

Featured stories

View more in Document understanding

Anthropic reports $42 billion loss in leaked IPO prospectus

Sep 29, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources

Story comments

Loading comments…

Related Projects

H Company

Topics

Document understandingAI foundation modelsRetrieval-augmented generation (RAG)Multimodal modelsAI research & benchmarks

Featured stories

View more in Document understanding

Anthropic reports $42 billion loss in leaked IPO prospectus

Sep 29, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources