AI Inference Silicon

Dedicated application-specific processors and accelerators engineered for high-throughput, low-latency AI model inference in data centers and edge deployments. Architectures overcome the memory wall by utilizing massive on-chip SRAM, high-bandwidth memory (HBM3e), and specialized matrix processing engines. Commercial implementations include Etched Sohu ASICs, Groq Language Processing Units (LPUs), Tenstorrent Wormhole, and Cerebras CS-3 systems.

Hidden Factory

AI Inference Silicon

1968 → 2026
Swipe roadmap
Exponential Industry Backdrop
milestone
1968
Intel founded; Grove first hire
IntelFairchild SemiconductorAndrew S. Grove
milestone
May 14, 2020
NVIDIA A100 Ampere GPU in full production (GTC 2020)
NVIDIA CorporationNVIDIA Hopper architecture and H100 announced (GTC 2022)Meta Llama 2 paper reports pretraining on NVIDIA A100-80GB
milestone
Sep 20, 2022
NVIDIA H100 Hopper GPU in full production
NVIDIA CorporationNVIDIA Hopper architecture and H100 announced (GTC 2022)Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs
milestone
Jul 23, 2024
Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs
MetaNVIDIA H100 Hopper GPU in full production
milestone
Oct 23, 2025
NVIDIA documents Unsloth QLoRA fine-tune on GeForce RTX 5090
NVIDIA CorporationGeForce RTX 5090 available at $1,999
milestone
Jun 29, 2026
Meta MTIA 300 presented at ISCA 2026
Meta
milestone
Aug 25, 2026
OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026
OpenAIBroadcom
2026
milestone

OpenAI Unpacks Jalapeño Custom LLM Inference ASIC at Hot Chips 2026

OpenAI presented architecture and benchmark results for Jalapeño at Hot Chips 2026, detailing its reticle-sized 64-slice inference ASIC with 216 GB HBM4, spatial programming model, and 9-month RTL-to-tapeout development cycle with Broadcom /OpenAI/ /Tom's Hardware/.

Related: OpenAI · organization Broadcom · organization
milestone

Ornn OCPI still prices A100, H100, H200, B200, and RTX 5090

Ornn Compute Price Index hourly API last_updated 2026-08-22T14:02:53Z: A100 SXM4 $1.06, H100 SXM $2.85, H200 $4.91, B200 $6.73, RTX 5090 $0.53 per GPU-hour. Settled daily index 2026-08-21T20:00:00Z puts RTX 5090 at $0.52. Free-tier docs list exactly these five GPUs. Snapshot of active rental-market financial impact — not a claim that any of them is still the frontier pretrain SKU /Ornn/.

Related: Ornn · organization
milestone

Etched delivers first inference rack to Jane Street

Etched PR Aug 18, 2026: first rack shipped last month; Jane Street said it tested the chip, is pleased with early results, and has its own rack running in its data center. Treated as the announcement day; ship month is July 2026 without a calendar day /GlobeNewswire/.

Related: Etched · organization
milestone

Meta MTIA 300 presented at ISCA 2026

MTIA Team (Meta Platforms) paper on MTIA 300, Meta's first training chip with built-in NIC chiplets and collective offloading engines, listed for ISCA 2026 (Raleigh, main program Jun 29–Jul 1). start_date is the ISCA 2026 main-program start, not a sourced individual talk slot /Meta Platforms / ISCA 2026/.

Related: Meta · organization
milestone

Vector Core fully disaggregated inference demonstration

Xeon 6 orchestration, SambaNova SN40 decode, NVIDIA Blackwell prefill from Vector Core LA data center; Together.ai first commercial customer.

Related: Vector Core Compute · organization Intel · organization SambaNova Systems · organization
milestone

US Patent Office Publishes OpenAI Patent US20260147536A1 on Hardware Alignment Accelerators

The USPTO published OpenAI's patent US20260147536A1 ('Alignment in Hardware Accelerators') describing data alignment mechanisms and mantissa scaling in compute-in-memory (CIM) macro arrays for high-efficiency neural accelerators /US Patent and Trademark Office/ /X/.

Related: OpenAI · organization
2025
milestone

NVIDIA documents Unsloth QLoRA fine-tune on GeForce RTX 5090

NVIDIA Technical Blog (Oct 23, 2025) reports Unsloth QLoRA fine-tuning on a GeForce RTX 5090 (32GB, Alpaca, Llama 3.1 8B-class table) and states a single Blackwell GPU can fine-tune models with as many as 40 billion parameters. This is the sourced 5090 training use-case — adapter fine-tune, not frontier pretrain. Explains why 5090 still clears a ~$0.52–0.53/hr Ornn spot without having trained Llama/GPT-class base models /NVIDIA Technical Blog/.

Related: NVIDIA Corporation · organization GeForce RTX 5090 available at $1,999 · process
milestone

GeForce RTX 5090 available at $1,999

CES PR (Jan 6, 2025) stated GeForce RTX 5090 would be available on Jan 30 at $1,999 (3,352 AI TOPS). Shipping day for the consumer Blackwell SKU that later appears on Ornn OCPI as the cheapest free-tier GPU /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization NVIDIA GeForce RTX 5090 announced at CES 2025 · process NVIDIA documents Unsloth QLoRA fine-tune on GeForce RTX 5090 · process
milestone

NVIDIA GeForce RTX 5090 announced at CES 2025

CES 2025 (Jan 6): NVIDIA unveiled GeForce RTX 50 Series on Blackwell. RTX 5090 is 92 billion transistors and over 3,352 AI TOPS; NVIDIA said it would be available Jan 30 at $1,999. Consumer SKU — local inference and creator AI, not a data-center HBM training GPU. No sourced SOTA foundation model was pretrained on RTX 5090 /NVIDIA Newsroom/ /NVIDIA GeForce News/.

Related: NVIDIA Corporation · organization NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024) · process GeForce RTX 5090 available at $1,999 · process
2024
milestone

Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs

Meta AI blog (Jul 23, 2024) introducing Llama 3.1, including 405B as the first frontier-level open-source Llama. Meta said it pushed training to over 16 thousand H100 GPUs on 15+ trillion tokens, making 405B the first Llama trained at that scale. Primary sourced breakthrough-model evidence for H100's Ornn price — the Hopper fleet that trained the 2024 open frontier model still rents in 2026 /Meta AI/.

Related: Meta · organization NVIDIA H100 Hopper GPU in full production · process
milestone

NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024)

GTC 2024 (Mar 18) launch of the NVIDIA Blackwell platform: Blackwell GPU architecture, GB200 Grace Blackwell Superchip, liquid-cooled GB200 NVL72 rack-scale system (72 Blackwell GPUs + 36 Grace CPUs on fifth-gen NVLink), HGX B200, sixth-generation air-cooled DGX B200 (8 Blackwell GPUs, two 5th Gen Intel Xeon; up to 144 PFLOPS FP4), and broad cloud/OEM adoption commitments. Foundational rackscale-AI milestone that established the NVL72 rack product line later continued by GB300 and Vera Rubin NVL72; DGX B200 is the traditional air-cooled 8-GPU SuperPOD/BasePOD building block, not NVL72 /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization Jensen Huang · person NVIDIA Vera Rubin platform announcement · process
2023
milestone

NVIDIA H200 HGX announced (SC23)

SC23 (Nov 13, 2023): NVIDIA announced HGX H200, the first GPU with HBM3e (141GB at 4.8 TB/s; nearly double A100 capacity and 2.4x bandwidth). NVIDIA said H200 would nearly double Llama 2 70B inference vs H100. Availability from 2Q 2024 as a compatible HGX H100 successor. Hopper memory-refresh branch, not a new architecture /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization NVIDIA Hopper architecture and H100 announced (GTC 2022) · process NVIDIA Blackwell platform and GB200 NVL72 announcement (GTC 2024) · process
milestone

Meta Llama 2 paper reports pretraining on NVIDIA A100-80GB

arXiv 2307.09288 (Jul 18, 2023): Meta released Llama 2 (7B–70B). Pretraining used NVIDIA A100s on RSC and internal production clusters; a cumulative 3.3M GPU hours on A100-80GB (TDP 400W or 350W). This is the sourced breakthrough-model evidence for A100's remaining financial impact — not that A100 still trains 2026 frontier runs, but that the open-weight Llama 2 generation ran on the Ampere fleet /arXiv/.

Related: Meta · organization NVIDIA A100 Ampere GPU in full production (GTC 2020) · process
2022
milestone

NVIDIA H100 Hopper GPU in full production

GTC Fall 2022 (Sep 20): NVIDIA announced H100 in full production, with partners planning an October rollout of Hopper-based products and services. DGX H100 (eight H100, 32 petaflops FP8) available to order. This is when Hopper became a rentable/shipped fleet, not only a GTC slide /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization NVIDIA Hopper architecture and H100 announced (GTC 2022) · process Meta Llama 3.1 405B trained on over 16 thousand H100 GPUs · process
milestone

NVIDIA H100 Transformer Engine blog / Hopper TE introduction coverage

NVIDIA Blog (Dave Salvator) details Transformer Engine on Hopper H100: mixed FP8/FP16 Tensor Core training/inference for transformers, MoE 395B and Megatron 530B speedups vs prior gen. Updated Aug 8, 2023 to note Ada Lovelace incorporation of Transformer Engine.

Related: NVIDIA Corporation · organization
milestone

NVIDIA Hopper architecture and H100 announced (GTC 2022)

GTC 2022 (Mar 22): NVIDIA announced Hopper, succeeding Ampere, and the first Hopper GPU H100 (80 billion transistors, Transformer Engine, HBM3, fourth-gen NVLink). Jensen Huang called H100 the engine of the world's AI infrastructure as data centers become AI factories. Availability stated as starting in the third quarter of 2022 /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization NVIDIA A100 Ampere GPU in full production (GTC 2020) · process NVIDIA H100 Hopper GPU in full production · process
2020
milestone

NVIDIA A100 Ampere GPU in full production (GTC 2020)

GTC 2020 (May 14): NVIDIA announced the first Ampere-architecture GPU, A100, in full production and shipping. More than 54 billion transistors on 7nm; unifies AI training and inference with up to 20x vs prior gen; MIG up to seven instances; third-gen NVLink. DGX A100 (eight A100s) announced the same day. Trunk of the modern NVIDIA LLM-training GPU line that Ornn still prices in 2026 /NVIDIA Newsroom/.

Related: NVIDIA Corporation · organization NVIDIA Hopper architecture and H100 announced (GTC 2022) · process Meta Llama 2 paper reports pretraining on NVIDIA A100-80GB · process
1987
milestone

Andy Grove becomes Intel CEO

Grove became president 1979, CEO 1987, chairman 1997–2005. Under his leadership Intel produced the 386 and Pentium; annual revenue rose from $1.9B to more than $26B (span not dated in the PR) /Intel/.

Related: Intel · organization Andrew S. Grove · person
1974
milestone

Intel operations in Oregon begin

Innovating and Investing in Oregon Since 1974 /Intel/.

Related: Intel · organization
1968
milestone

Intel founded; Grove first hire

Noyce and Moore left Fairchild to found Intel in 1968; Grove was their first hire /Intel/.

Related: Intel · organization Fairchild Semiconductor · organization Andrew S. Grove · person