Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
H Company releases NeoMME multimodal encoders for visual document retrieval
00

H Company releases NeoMME multimodal encoders for visual document retrieval

Sep 6, 2026

H Company has released NeoMME, a family of 260M and 800M single-tower multimodal encoders designed to optimize visual document retrieval by dropping the traditional vision tower and causal decoder. Pretrained using discrete masked diffusion, the NeoMME-260M model achieves 0.523 nDCG@10 on ViDoRe v3, matching the performance of the 14.4x larger ColQwen2.5 model. While text-only retrieval remains a weak spot, NeoMME offers high throughput and efficient index compression down to 6 kB per page.

NeoMME multimodal encoder architecture

  • ▪The NeoMME models use an ALBERT-style factorized embedding for text and a 2-layer MLP to project non-overlapping 32x32 image patches, with no patch-merging module or SigLIP2 tower.
  • ▪NeoMME models support a 16,384-token context window, which is sufficient to process two standard 3,840x2,160 4K UHD images after patching.
  • ▪The NeoMME architecture drops both the separately pretrained vision tower and the causal decoder typically found in generative vision-language models to reduce parameter and compute overhead.
  • ▪H Company released NeoMME, a family of 260M and 800M single-tower, bidirectional multimodal encoders that process multilingual text tokens and raw 32x32 RGB image patches through the same Transformer layers.
  • ▪The exact parameter counts for the NeoMME models are 262,937,906 for the 260M model and 793,715,032 for the 800M model, and they ship under the Apache 2.0 license.

Masked diffusion pretraining methodology

  • ▪During NeoMME pretraining, text-only segments draw a corruption rate from 0 to 1, while multimodal segments draw from 0.30 to 1 to force the model to read the page.
  • ▪A cross-modal ablation probe showed that at 90% masking, visible page patches raised masked-token accuracy by 38.4 points for the NeoMME-260M model and 40.5 points for the NeoMME-800M model.
  • ▪NeoMME is pretrained using discrete masked diffusion over text, which is optionally conditioned on visible image patches.
  • ▪The pretraining of NeoMME processed approximately 524 billion packed input tokens (including 290 billion text-only tokens) using 16 and 32 H100 accelerators for the 260M and 800M models respectively.

Visual document retrieval performance

  • ▪The NeoMME-Retriever-800M model scores 0.556 nDCG@10 on the ViDoRe v3 benchmark, landing 0.9 points behind the similarly sized Vultron Retriever Flash.
  • ▪On the ViDoRe v1 and v2 benchmarks, the NeoMME models achieve scores of 0.860/0.522 (for the 260M model) and 0.874/0.559 (for the 800M model) nDCG@5.
  • ▪The NeoMME-Retriever-260M model scores 0.523 nDCG@10 on the ViDoRe v3 benchmark, which is within 0.002 of the 3.75B-parameter ColQwen2.5-v0.2 model.

Index compression techniques

  • ▪Applying hierarchical token pooling at factor 10 with int8 queries and documents reduces the NeoMME index size to 39.0 kB per page (a 39.4x reduction) while retaining 99.16% of baseline nDCG@10.
  • ▪Late-interaction indexes for NeoMME yield 4,162 vectors (about 1.5 MB in float32) per 2048x2048 page on the ViDoRe v3 benchmark.

Throughput benchmarks

  • ▪The NeoMME-260M model indexes 51.3 pages per second on a single NVIDIA L40S GPU at a matched 2048x2048 input, which is 1.97x faster than ColModernVBERT's 26.0 pages per second.
  • ▪The NeoMME-260M model encodes a query in 78.3 ms on a CPU-only host.

Text-only retrieval limitations

  • ▪The weaker text retrieval performance of NeoMME is attributed to supervision scale, as it was trained on approximately 430K text query examples compared to mLateOn's 660M contrastive examples.
  • ▪On the BEIR-15 text retrieval benchmark, NeoMME's late interaction reaches 0.4881 (260M) and 0.5126 (800M), compared to 0.5722 for the 149M-parameter LateOn model.

1 source

Marktechpost
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
View source article

Featured stories

View more in Document understanding

Anthropic reports $42 billion loss in leaked IPO prospectus

Sep 29, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources

Story comments

Loading comments…

Related Projects

H Company

Topics

Document understandingRetrieval-augmented generation (RAG)AI foundation modelsAI research & benchmarksMultimodal models

Featured stories

View more in Document understanding

Anthropic reports $42 billion loss in leaked IPO prospectus

Sep 29, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources