AI Inference Silicon
- Exponential Industry
- Technology Area
- 20 Milestones
Dedicated application-specific processors and accelerators engineered for high-throughput, low-latency AI model inference in data centers and edge deployments. Architectures overcome the memory wall by utilizing massive on-chip SRAM, high-bandwidth memory (HBM3e), and specialized matrix processing engines. Commercial implementations include Etched Sohu ASICs, Groq Language Processing Units (LPUs), Tenstorrent Wormhole, and Cerebras CS-3 systems.
Hidden Factory
AI Inference Silicon
OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026
OpenAI presented architecture and benchmark results for Jalapeño at Hot Chips 2026, detailing its reticle-sized 64-slice inference ASIC with 216 GB HBM4, spatial programming model, and 9-month RTL-to-tapeout development cycle with Broadcom /OpenAI/ /Tom's Hardware/.
Ornn OCPI still prices A100, H100, H200, B200, and RTX 5090
Ornn Compute Price Index hourly API last_updated 2026-08-22T14:02:53Z: A100 SXM4 $1.06, H100 SXM $2.85, H200 $4.91, B200 $6.73, RTX 5090 $0.53 per GPU-hour. Settled daily index 2026-08-21T20:00:00Z puts RTX 5090 at $0.52. Free-tier docs list exactly these five GPUs. Snapshot of active rental-market financial impact — not a claim that any of them is still the frontier pretrain SKU /Ornn/.
Etched delivers first inference rack to Jane Street
Etched PR Aug 18, 2026: first rack shipped last month; Jane Street said it tested the chip, is pleased with early results, and has its own rack running in its data center. Treated as the announcement day; ship month is July 2026 without a calendar day /GlobeNewswire/.
Meta MTIA 300 presented at ISCA 2026
MTIA Team (Meta Platforms) paper on MTIA 300, Meta's first training chip with built-in NIC chiplets and collective offloading engines, listed for ISCA 2026 (Raleigh, main program Jun 29–Jul 1). start_date is the ISCA 2026 main-program start, not a sourced individual talk slot /Meta Platforms / ISCA 2026/.
Vector Core fully disaggregated inference demonstration
Xeon 6 orchestration, SambaNova SN40 decode, NVIDIA Blackwell prefill from Vector Core LA data center; Together.ai first commercial customer.
US Patent Office Publishes OpenAI Patent US20260147536A1 on Hardware Alignment Accelerators
The USPTO published OpenAI's patent US20260147536A1 ('Alignment in Hardware Accelerators') describing data alignment mechanisms and mantissa scaling in compute-in-memory (CIM) macro arrays for high-efficiency neural accelerators /US Patent and Trademark Office/ /X/.
NVIDIA documents Unsloth QLoRA fine-tune on GeForce RTX 5090
NVIDIA Technical Blog (Oct 23, 2025) reports Unsloth QLoRA fine-tuning on a GeForce RTX 5090 (32GB, Alpaca, Llama 3.1 8B-class table) and states a single Blackwell GPU can fine-tune models with as many as 40 billion parameters. This is the sourced 5090 training use-case — adapter fine-tune, not frontier pretrain. Explains why 5090 still clears a ~$0.52–0.53/hr Ornn spot without having trained Llama/GPT-class base models /NVIDIA Technical Blog/.
GeForce RTX 5090 available at $1,999
CES PR (Jan 6, 2025) stated GeForce RTX 5090 would be available on Jan 30 at $1,999 (3,352 AI TOPS). Shipping day for the consumer Blackwell SKU that later appears on Ornn OCPI as the cheapest free-tier GPU /NVIDIA Newsroom/.
NVIDIA GeForce RTX 5090 announced at CES 2025
CES 2025 (Jan 6): NVIDIA unveiled GeForce RTX 50 Series on Blackwell. RTX 5090 is 92 billion transistors and over 3,352 AI TOPS; NVIDIA said it would be available Jan 30 at $1,999. Consumer SKU — local inference and creator AI, not a data-center HBM training GPU. No sourced SOTA foundation model was pretrained on RTX 5090 /NVIDIA Newsroom/ /NVIDIA GeForce News/.
Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs
Meta AI blog (Jul 23, 2024) introducing Llama 3.1, including 405B as the first frontier-level open-source Llama. Meta said it pushed training to over 16 thousand H100 GPUs on 15+ trillion tokens, making 405B the first Llama trained at that scale. Primary sourced breakthrough-model evidence for H100's Ornn price — the Hopper fleet that trained the 2024 open frontier model still rents in 2026 /Meta AI/.
NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024)
GTC 2024 (Mar 18) launch of the NVIDIA Blackwell platform: Blackwell GPU architecture, GB200 Grace Blackwell Superchip, liquid-cooled GB200 NVL72 rack-scale system (72 Blackwell GPUs + 36 Grace CPUs on fifth-gen NVLink), HGX B200, sixth-generation air-cooled DGX B200 (8 Blackwell GPUs, two 5th Gen Intel Xeon; up to 144 PFLOPS FP4), and broad cloud/OEM adoption commitments. Foundational rackscale-AI milestone that established the NVL72 rack product line later continued by GB300 and Vera Rubin NVL72; DGX B200 is the traditional air-cooled 8-GPU SuperPOD/BasePOD building block, not NVL72 /NVIDIA Newsroom/.
NVIDIA H200 HGX announced (SC23)
SC23 (Nov 13, 2023): NVIDIA announced HGX H200, the first GPU with HBM3e (141GB at 4.8 TB/s; nearly double A100 capacity and 2.4x bandwidth). NVIDIA said H200 would nearly double Llama 2 70B inference vs H100. Availability from 2Q 2024 as a compatible HGX H100 successor. Hopper memory-refresh branch, not a new architecture /NVIDIA Newsroom/.
Meta Llama 2 paper reports pretraining on NVIDIA A100-80GB
arXiv 2307.09288 (Jul 18, 2023): Meta released Llama 2 (7B–70B). Pretraining used NVIDIA A100s on RSC and internal production clusters; a cumulative 3.3M GPU hours on A100-80GB (TDP 400W or 350W). This is the sourced breakthrough-model evidence for A100's remaining financial impact — not that A100 still trains 2026 frontier runs, but that the open-weight Llama 2 generation ran on the Ampere fleet /arXiv/.
NVIDIA H100 Hopper GPU in full production
GTC Fall 2022 (Sep 20): NVIDIA announced H100 in full production, with partners planning an October rollout of Hopper-based products and services. DGX H100 (eight H100, 32 petaflops FP8) available to order. This is when Hopper became a rentable/shipped fleet, not only a GTC slide /NVIDIA Newsroom/.
NVIDIA H100 Transformer Engine blog / Hopper TE introduction coverage
NVIDIA Blog (Dave Salvator) details Transformer Engine on Hopper H100: mixed FP8/FP16 Tensor Core training/inference for transformers, MoE 395B and Megatron 530B speedups vs prior gen. Updated Aug 8, 2023 to note Ada Lovelace incorporation of Transformer Engine.
NVIDIA Hopper architecture and H100 announced (GTC 2022)
GTC 2022 (Mar 22): NVIDIA announced Hopper, succeeding Ampere, and the first Hopper GPU H100 (80 billion transistors, Transformer Engine, HBM3, fourth-gen NVLink). Jensen Huang called H100 the engine of the world's AI infrastructure as data centers become AI factories. Availability stated as starting in the third quarter of 2022 /NVIDIA Newsroom/.
NVIDIA A100 Ampere GPU in full production (GTC 2020)
GTC 2020 (May 14): NVIDIA announced the first Ampere-architecture GPU, A100, in full production and shipping. More than 54 billion transistors on 7nm; unifies AI training and inference with up to 20x vs prior gen; MIG up to seven instances; third-gen NVLink. DGX A100 (eight A100s) announced the same day. Trunk of the modern NVIDIA LLM-training GPU line that Ornn still prices in 2026 /NVIDIA Newsroom/.
Andy Grove becomes Intel CEO
Grove became president 1979, CEO 1987, chairman 1997–2005. Under his leadership Intel produced the 386 and Pentium; annual revenue rose from $1.9B to more than $26B (span not dated in the PR) /Intel/.
Intel operations in Oregon begin
Innovating and Investing in Oregon Since 1974 /Intel/.
Intel founded; Grove first hire
Noyce and Moore left Fairchild to found Intel in 1968; Grove was their first hire /Intel/.