Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

AI inference (scaling) stories

Sep 16, 2026

Apple developing enterprise AI server with M8 Ultra chips for 2029 launch

Apple is working on an enterprise server using two or four M8 Ultra chips for the AI inference market, with discussions underway with Nvidia about networking equipment, according to The Information. Launch expected no earlier than 2029.

Sep 16, 2026·4 sources
00
Sep 15, 2026

Crusoe signs multiyear cloud deal with Perplexity AI for Nvidia GB300 chip access

Data center operator Crusoe has secured a multiyear cloud capacity agreement with AI search company Perplexity AI, providing access to Nvidia GB300 processors to support the startup's model training and inference workloads.

Sep 15, 2026·3 sources
00

European AI chip startups Axelera and Euclyd raise funding, challenge Nvidia in inference market

Dutch startup Axelera AI launched its second-generation Europa chip and secured AI factory supply contracts with over 600 customers, while Euclyd raised €200M in Series A funding co-led by Samsung to develop its CRAFTWERK inference chip as an alternative to GPU-based solutions.

Sep 15, 2026·4 sources
00
Sep 10, 2026

Positron AI raises $875 million at $5 billion valuation as AI chip demand surges

AI chip startup Positron AI has raised $875 million in its latest funding round, more than quadrupling its valuation to $5 billion in seven months. The company develops chips designed to run AI models more quickly.

Sep 10, 2026·2 sources
00

DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage

Chinese AI company DeepSeek launched its V4.1-Flash model on September 10, 2026, featuring a 552-billion-parameter architecture that cuts memory requirements by 75% and offers inference pricing as low as $0.003 per million tokens. The model claims to match or exceed performance of competitors like GPT-6 Astra, Claude Opus 5, and GPT-5.6 Sol on various benchmarks while operating at a fraction of the cost.

Sep 10, 2026·12 sources
00

DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage

Chinese AI company DeepSeek launched its V4.1-Flash model on September 10, 2026, featuring a 552-billion-parameter architecture that cuts memory requirements by 75% and offers inference pricing as low as $0.003 per million tokens. The model claims to match or exceed performance of competitors like GPT-6 Astra, Claude Opus 5, and GPT-5.6 Sol on various benchmarks while operating at a fraction of the cost.

Sep 10, 2026·12 sources
00
Sep 3, 2026

Nvidia launches PAIR tool for home AI computing

Nvidia announced Personal AI Router (PAIR), a free tool that syncs home computers to create a distributed network for local AI inference tasks, turning household devices into a mini data center.

Sep 3, 2026·3 sources
00
Sep 1, 2026

Dell forecasts AI inference token demand to surge 3,400% by 2030

Dell Technologies projected at its Dell Technologies World event that global AI token consumption will increase 3,400% by 2030, while booking a record $60.9 billion in AI server orders during Q2 fiscal 2027.

Sep 1, 2026·3 sources
00

Perplexity open sources Lily inference engine and launches Hybrid Compute for local AI processing

Perplexity has launched Hybrid Compute, allowing users to combine cloud and local AI models while keeping sensitive data on their Mac, and open sourced Lily, a Rust + Metal inference engine for running Qwen3.6-35B-A3B on Apple Silicon that powers the local processing capability.

Sep 1, 2026·2 sources
00
Aug 28, 2026

Tencent releases Hy4 preview, open-source AI model with self-improvement capabilities

Tencent has unveiled Hy4 Preview, a 770-billion-parameter open-source Mixture-of-Experts model that uses a recursive self-improvement loop to boost inference throughput by 31.8%. The release, trained on Tencent's vast product ecosystem, positions the company back in the top tier of China's AI competition.

Aug 28, 2026·3 sources
00
Aug 25, 2026

OpenAI unveils Jalapeño chip with performance claims exceeding Nvidia GB300

OpenAI published benchmark results for its custom Jalapeño inference chip built with Broadcom, claiming it delivers more AI work per watt and faster responses than Nvidia's GB300 processors with a 50% cost advantage.

Aug 25, 2026·2 sources
00

OpenAI unveils Jalapeño chip with performance gains over Nvidia flagship GPU

OpenAI presented benchmarks at Hot Chips conference showing its 700W Jalapeño ASIC, co-developed with Broadcom, delivers up to 1.9x throughput per kilowatt and 3.6x lower latency compared to Nvidia's 1,400W flagship GPU. The chip is designed for fast AI inference at scale.

Aug 25, 2026·4 sources
00
Aug 24, 2026

Nvidia begins mass production of Groq 3 LPX AI inference chips

Nvidia has put its Groq 3 LPX racks into full production following its $20 billion acquisition of Groq, with the SRAM-based language processing units achieving 3,400 tokens per second and Samsung manufacturing the 4-nanometer chips. The first deployment is set for Nebius Token Factory before year-end.

Aug 24, 2026·4 sources
00

d-Matrix unveils Raptor AI accelerator with 100 TB/s bandwidth using stacked DRAM architecture

d-Matrix presented Raptor at Hot Chips 2026, featuring a TSMC 4nm compute die bonded face-to-face at 36-micron pitch directly onto custom DRAM, achieving 100 terabytes per second memory bandwidth while using one-sixth the energy per bit of HBM3.

Aug 24, 2026·2 sources
00
Aug 21, 2026

Starcloud raises $250 million for orbital AI data centers at $2.3 billion valuation

Starcloud, a startup developing satellites that can perform AI inference in orbit, has raised $250 million in an extension to its March Series A funding round, bringing its valuation to $2.3 billion. The company is building data centers in space as launch access becomes increasingly constrained.

Aug 21, 2026·4 sources
00
Aug 18, 2026

NVIDIA releases TensorRT Model Connect in public preview for streamlined AI inference

NVIDIA has launched TensorRT Model Connect (TRTMC) as an open-source public preview tool that enables developers to convert Hugging Face or local checkpoints to TensorRT inference using just two commands, simplifying the deployment of AI models.

Aug 18, 2026·1 source
00
Aug 17, 2026

Groq raises $350M at $3.5B valuation as it pivots from AI chips to Nvidia-powered cloud infrastructure

AI infrastructure startup Groq has raised $350 million in Series A funding at a $3.5 billion valuation, marking a strategic pivot from developing its own AI chips to becoming a neocloud provider offering GPU and AI infrastructure services. Notably, Nvidia has joined the funding round as the company shifts to using Nvidia-powered infrastructure.

Aug 17, 2026·3 sources
00
Aug 14, 2026

French startup Kog pursues AI inference optimization amid GPU chip competition

French startup Kog is developing technology to extract more AI inference performance from existing GPUs, entering a competitive market where purpose-built chip makers like Cerebras recently debuted successfully with its May IPO.

Aug 14, 2026·1 source
00
Aug 13, 2026

OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering 14x faster speeds via Cerebras partnership

OpenAI introduced Ultrafast, a new inference tier that runs GPT-5.6 Sol at up to 750 tokens per second—14 times faster than Standard mode—powered by Cerebras wafer-scale chips from their $10 billion partnership. The service creates a three-tier pricing structure targeting real-time applications like voice AI, financial research, and incident response.

Aug 13, 2026·4 sources
00
Aug 11, 2026

IBM and Together AI sign $240 million deal for Nvidia-powered AI inference cluster

IBM and startup Together AI announced a multi-year $240 million agreement to build a large-scale artificial intelligence inference cluster on IBM Cloud using Nvidia systems. The deployment will feature Nvidia's HGX B300 systems and is scheduled to launch in Q1 2027.

Aug 11, 2026·5 sources
00

IBM and Together AI sign $240 million deal to build Nvidia-powered AI inference cluster

IBM and startup Together AI announced a multi-year $240 million agreement to build a large-scale artificial intelligence inference cluster on IBM Cloud using Nvidia systems. The deal reflects IBM's bet that enterprise AI inference—the process of running trained models at scale—represents a major growth opportunity.

Aug 11, 2026·3 sources
00
Aug 6, 2026

AMD Acquires AI Chip Startup Taalas to Boost Inference Performance with Model-in-Silicon Technology

AMD announced the acquisition of Toronto-based AI chip startup Taalas for an undisclosed sum on August 6, 2026. Taalas specializes in etching AI model weights directly into silicon to enhance inference performance, part of AMD's strategy to challenge Nvidia's dominance in AI hardware.

Aug 6, 2026·6 sources
00

AMD Acquires AI Chip Startup Taalas to Boost Inference Performance with Model-in-Silicon Technology

AMD announced the acquisition of Toronto-based AI chip startup Taalas for an undisclosed sum on August 6, 2026. Taalas specializes in etching AI model weights directly into silicon to enhance inference performance, part of AMD's strategy to challenge Nvidia's dominance in AI hardware.

Aug 6, 2026·5 sources
00
Aug 3, 2026

Olix Computing raises $312 million for optical AI inference hardware

AI hardware startup Olix Computing announced a $312 million funding round backed by Netflix co-founder to develop optical inference appliances for AI workloads.

Aug 3, 2026·1 source
00
Jul 23, 2026

VAST Data Expands Partnership with AMD to Integrate AI Operating System with EPYC CPUs and Instinct GPUs

VAST Data announced an expanded collaboration with AMD to integrate its AI Operating System with AMD's 6th Gen EPYC CPUs, Instinct GPUs, and Pensando AI NICs, targeting AI cloud providers and enterprises building high-performance AI infrastructure for inference workloads.

Jul 23, 2026·5 sources
00
Jul 17, 2026

General Compute lands $400 million loan for AI inference chips

AI inference cloud startup General Compute secured a $400 million loan from Upper90, reportedly the first deal to use inference-specific chips as collateral.

Jul 17, 2026·1 source
00
Jul 8, 2026

Cognition Launches SWE-1.7 AI Model, Claims Near-Frontier Performance at Lower Cost

Cognition released SWE-1.7, its most capable AI model to date, which the company says performs within a few points of leading frontier models while offering significantly reduced costs and 1000 tokens per second speed. The company reports continued gains from scaling reinforcement learning.

Jul 8, 2026·3 sources
00
Jul 7, 2026

DeepSeek Developing Own AI Chip to Reduce Reliance on Nvidia and Huawei

Chinese AI startup DeepSeek is developing its own chip designed for inference workloads, according to three sources familiar with the matter, in a move to reduce dependence on Nvidia and Huawei chips.

Jul 7, 2026·3 sources
00
Jun 30, 2026

AI Chip Startup Etched Reaches $5B Valuation with $1B in Orders and $800M Funding Round

Etched, a startup developing specialized AI inference chips for Transformer models, announced it has secured $1 billion in contract orders and raised $800 million in funding, reaching a $5 billion valuation after TSMC successfully manufactured its Sohu chip.

Jun 30, 2026·12 sources
00

OpenAI Achieves 50% Reduction in AI Inference Costs Through New Optimization Technique

OpenAI has reportedly developed a system-level optimization method that cuts the inference costs of running its AI models by approximately 50%, a breakthrough that could significantly improve the economics of operating large language models at scale.

Jun 30, 2026·8 sources
00

AI chip startup Etched raises $800M, emerges from stealth with working inference chip and $1B in customer contracts

Etched, a company designing chips specifically for running AI models, has raised $800 million from investors including Jane Street and TSMC-linked VentureTech Alliance, while announcing it has secured over $1 billion in signed customer contracts.

Jun 30, 2026·2 sources
00
Jun 24, 2026

OpenAI unveils Jalapeño, its first custom AI chip developed with Broadcom

OpenAI announced its first custom-built AI inference processor called Jalapeño, designed in partnership with Broadcom and optimized for running large language models. The chip is expected to be deployed at scale by late 2026 to power ChatGPT and other AI products.

Jun 24, 2026·9 sources
00
Jun 10, 2026

Google releases DiffusionGemma, an experimental AI model with 4x faster text generation

Google has launched DiffusionGemma, an open experimental model that delivers up to 4x faster inference on dedicated GPUs, targeting speed-critical and interactive local applications.

Jun 10, 2026·1 source
00
May 9, 2026

LightSeek Foundation Releases TokenSpeed Open-Source Inference Engine

LightSeek Foundation launched TokenSpeed, an open-source LLM inference engine targeting TensorRT-LLM-level performance specifically optimized for agentic coding workloads like Claude Code, Codex, and Cursor.

May 9, 2026·1 source
00
May 7, 2026

Google Releases Multi-Token Prediction Drafters for Gemma 4 with 3x Speedup

Google AI released Multi-Token Prediction drafters for the Gemma 4 model family using speculative decoding, achieving up to 3x faster inference without quality loss by addressing memory-bandwidth bottlenecks.

May 7, 2026·2 sources
00
Apr 12, 2026

NVIDIA Releases AITune Open-Source Inference Toolkit for Automatic Backend Optimization

NVIDIA launched AITune, an open-source inference toolkit that automatically identifies the fastest inference backend for any PyTorch model, addressing the deployment gap between research and production.

Apr 12, 2026·1 source
00
Mar 30, 2026

South Korean AI Chip Startup Rebellions Raises $400M at $2.3B Valuation in Pre-IPO Round

Rebellions, which designs chips specifically for AI inference, raised $400 million at a $2.3 billion valuation ahead of a planned IPO later this year, positioning itself as another challenger to Nvidia's dominance.

Mar 30, 2026·1 source
00
Mar 28, 2026

Google Introduces TurboQuant Algorithm Reducing LLM Memory Usage by 6x

Google released TurboQuant, a compression algorithm that reduces large language model memory usage by at least 6x while improving performance, targeting AI inference cost reduction. The announcement caused significant market impact, with memory chip maker stocks reportedly declining.

Mar 28, 2026·2 sources
00

Top claims

  • ▪Anthropic has rented Mac Minis from Amazon Web Services to support its artificial intelligence workloads
  • ▪The expected 2029 release of Apple's planned enterprise AI server, powered by M8 Ultra chips, would mark Apple's first server product to hit the market in nearly two decades
  • ▪Apple's planned enterprise AI server, targeted for a 2029 launch, will be offered in two configurations, featuring either two or four of Apple's future M8 Ultra chips

Subtopics

Compute, chips & AI infrastructure32AI startups18AI tools & products11Large language models (LLMs)8Open-source AI7AI foundation models5AI research & benchmarks4AI for developers3Business & enterprise AI3Edge AI3Local AI3AI antitrust & competition2AI coding assistants2AI tokens2China tech & industrial strategy2Cloud Computing2Venture capital funding2AI accelerators1AI agents1AI data centers1AI hardware1AI infrastructure1AI privacy & surveillance1AI scaling laws1

Related timelines

AI Data Center Gold Rush

101 stories

Congress

108 stories

Crypto hacks

100 stories

Ebola outbreak

58 stories

Iran War

209 stories