H Company has released NeoMME, a family of 260M and 800M single-tower bidirectional multimodal encoders designed to optimize visual document retrieval. By eliminating separate vision towers and causal decoders, NeoMME processes text and raw image patches within a single Transformer. The 260M model achieves 0.523 nDCG@10 on ViDoRe v3, matching the performance of much larger models while offering high throughput and compression options.
NeoMME bidirectional encoder architecture
- ▪The NeoMME tokenizer is a whitespace-unconstrained BPE with a 131,072-entry vocabulary that emits 44.4% fewer tokens than ModernBERT across 14 target languages in FLORES-200 devtest.
- ▪H Company released NeoMME, a family of 260M and 800M parameter single-tower bidirectional multimodal encoders that process text and raw image patches through a single Transformer.
- ▪Both NeoMME models support a 16,384-token context window and utilize symmetric sliding-window attention on most layers, with global attention on every sixth layer and the final layer.
- ▪The NeoMME architecture processes text via an ALBERT-style factorized embedding and raw 32x32 RGB image patches via a 2-layer MLP, omitting a separate vision tower and causal decoder.
Masked diffusion pretraining approach
- ▪A cross-modal ablation probe showed that at 90% masking, visible page patches raised masked-token accuracy by 38.4 points for the NeoMME-260M model and 40.5 points for the NeoMME-800M model.
- ▪NeoMME models are pretrained using discrete masked diffusion over text, with multimodal segments drawing a corruption rate of 0.30 to 1 to force the model to read the page.
- ▪Pretraining of the NeoMME models processed approximately 524 billion packed input tokens, including roughly 290 billion text-only tokens, on NVIDIA H100 accelerators.
Visual document retrieval performance
- ▪On the ViDoRe v3 benchmark, the NeoMME-Retriever-260M model scored 0.523 nDCG@10, which is within 0.002 of the 3.75B-parameter ColQwen2.5-v0.2 model.
- ▪On the ViDoRe v3 benchmark, the NeoMME-Retriever-800M model scored 0.556 nDCG@10, landing 0.9 points behind the similarly sized Vultron Retriever Flash.
- ▪On ViDoRe v1 and v2 benchmarks, the NeoMME models achieved nDCG@5 scores of 0.860 and 0.522 for the 260M model, and 0.874 and 0.559 for the 800M model.
Text retrieval limitations
- ▪The authors of NeoMME attributed its weaker text retrieval performance to supervision scale, as NeoMME was trained on 430,000 text query examples compared to mLateOn's 660 million.
- ▪On the BEIR-15 text retrieval benchmark, NeoMME late-interaction scored 0.4881 for the 260M model and 0.5126 for the 800M model, compared to 0.5722 for the 149M-parameter LateOn model.
Index compression techniques
- ▪Applying hierarchical token pooling at factor 10 with int8 queries and documents reduced the NeoMME index size to 39.0 kB per page while retaining 99.16% of baseline nDCG@10.
- ▪Standard late-interaction indexing for NeoMME yields approximately 1.5 MB per ViDoRe v3 document page in float32 format.
Encoding throughput benchmarks
- ▪The NeoMME-260M model indexes 51.3 pages per second on a single NVIDIA L40S accelerator at a matched 2048x2048 input, which is 1.97 times faster than ColModernVBERT.
- ▪The NeoMME-260M model encodes a query in 78.3 milliseconds on a CPU-only host.
Story comments
Loading comments…