Subscribe
Sign in
Home
Main Site
Models & Research
Die Yield Calculator
Compliance
Archive
About
Latest
Top
Discussions
TPU Inference Externalization Full Steam Ahead - InferenceX
InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat
Sep 7
•
Alec Ibarra
,
Cam Quilici
,
Bryan Shan
,
Wenyao Gao
,
Daniel Nishball
,
Zane Fong
, and
Dylan Patel
108
9
Korea’s Trillion-Dollar Sovereign AI Investment: Nvidia Wins, Hynix Loses
Korea hosts a Squid Games, National AI Tournament, the best non-Chinese open source model gets eliminated, why Nvidia needs open source, implications…
Sep 1
•
Max Kan
,
Ray Wang
,
Myron Xie
, and
Dylan Patel
171
1
8
August 2026
Most Neoclouds Suck At Security
OpenAI vs HuggingFace, Container Escapes, Kernel Bypass, Network Policies, Security Keys, Multi-tenant Grafana, and a ClusterMAX 3.0 Preview
Aug 30
•
Jordan Nanos
,
Sam Harshe
,
Pratt Bhatt
,
Billy Cao
,
Jack Carson
, and
Dylan Patel
137
5
7
OpenAI Jalapeño: Better Than Nvidia Blackwell
OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets
Aug 25
•
Bryan Shan
,
Myron Xie
,
Jordan Nanos
,
Wega Chu
,
Clara Ee
, and
Dylan Patel
269
2
30
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200
Aug 24
•
Cam Quilici
,
Bryan Shan
,
Alec Ibarra
,
Daniel Nishball
,
Zane Fong
,
Kimbo Chen
, and
Dylan Patel
86
1
8
Are Open Models Catching Up?
Comparing open vs. closed models across the eras of frontier models, Is the gap narrowing?
Aug 21
•
Evan Cloutier
,
Max Kan
,
Jordan Nanos
, and
Dylan Patel
245
4
27
Cerebras's Next Generation CS-4: Fast Just Got Faster
Double the Performance, Double the Power, Double the Fun
Aug 19
•
Myron Xie
,
Bryan Shan
,
Wega Chu
,
Jordan Nanos
, and
Dylan Patel
138
1
6
$12B of US ratepayers' money wasted on a modeling mistake and PJM wants to do it again
American Grid design needs an overhaul, Why it is good to be full of cold air.
Aug 16
•
Robert Boswall
,
Reyk Knuhtsen
,
Jeremie Eliahou Ontiveros
,
Oliver Kennon
,
Dylan Patel
, and
Dashboard American
119
4
9
Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX
Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, high throughput prefill engine, high…
Aug 10
•
Bryan Shan
,
Daniel Nishball
,
Cam Quilici
,
Kimbo Chen
,
Alec Ibarra
, and
Dylan Patel
147
1
6
SpaceX 10GW in 2027 – Why It’s Real, Will Drive $300B ARR for SpaceX, and Why Microsoft Will Be the Largest Offtaker
Inference at 100B/GW/year, SpaceX's stellar pace, Microsoft's 10GW 2026 Awakening, Azure Can Grow Triple-Digits
Aug 7
•
Jeremie Eliahou Ontiveros
,
Reyk Knuhtsen
,
Jordan Nanos
,
Max Kan
,
Dylan Patel
, and
Muhammad Zuhair
299
11
20
Gemini is Cooked but GCP is Cooking
GCP YoY rev growth >100%, DeepMind's long term failure is Google Cloud's short term gain
Aug 7
•
Max Kan
,
Joey Brookhart
,
Doug O'Laughlin
, and
Dylan Patel
321
9
30
Kimi K3, The Manos, The Mythos, The Legendos
Kimi K3’s architecture: compressed memory, attention across depth, latent expert routing, and serving performance
Aug 3
•
Kimbo Chen
,
Shubham Choudhari
,
Bryan Shan
, and
Dylan Patel
175
12
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts