← Writing

LLM Year in Review

Hits

Foundation Model: Open vs Closed

Related post · I wrote a dedicated deep dive on this debate: Open vs Closed.

2025 has been a year of shifts in the foundation model landscape. While we've seen more powerful models from Google DeepMind, Anthropic and OpenAI, the shift between open source and closed source continues to evolve. Companies like Meta and xAI switched their previously open-source mode to closed source. OpenAI shipped gpt-oss (their 1st open-source model family post GPT-3) in August 2025.

Despite these industry shifts, my favorite AI breakthrough this year remains DeepSeek. DeepSeek proved that open source is the way. Their unwavering commitment to openness, even as giants like Meta retreated, has inspired a new generation of open-source models including Alibaba Qwen, Kimi, Xiaomi MiMo, MiniMax and AI2's OLMo. Open source enables innovation, research, and democratization of AI technology.

September 10, 2026
DeepSeek released DeepSeek-V4.1-Flash — MIT license, 1M context, $0.15/1M tokens (open weights)
September 3, 2026
OpenAI released GPT-6 Astra — first model to hit Critical-tier cybersecurity capability; Brockman calls it a step toward AGI (closed source)
September 1, 2026
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 — same model, different safeguards; Mythos restricted to trusted cybersecurity/bio access (closed source)
August 14, 2026
Z.ai released GLM-5.3, post-training only on the same 743B base. Against Fable 5: level on Terminal Bench 2.1 (88.2 vs 88.0), ahead on CyberGym (84.5 vs 83.8), behind on DeepSWE (66.9 vs 69.7) and ExploitBench (54.4 vs 78.0). Open weights two weeks out, after safety evaluation
August 12, 2026
Two releases in one day, each measured against Anthropic's Claude Fable 5 — 62 on the Artificial Analysis Intelligence Index, $7.70 per 1M blended tokens:
  • SpaceXAI's Grok 4.6 — one point off Fable 5 at 61, for $1.35 per 1M: ~82% cheaper (closed weights)
  • DeepSeek's DeepSeek-V4-Pro-0813 — GA build of the 1.6T/49B MoE, 1M context, for $0.18 per 1M; DeepSeek's own agent evals put it level with Fable 5 on Terminal Bench 2.1 (87.9 vs 88.0) and Cybergym (83.3 vs 83.1), behind on DeepSWE (62.7 vs 70.0) — open weights
August 2, 2026
Alibaba announced Qwen3.8-Max — their most capable model to date at 2.4T parameters, with open weights for both Qwen3.8-Max and Qwen3.8-27B coming the following week
July 16, 2026
Moonshot AI released Kimi K3: Open Frontier Intelligence — open weights at the frontier
January 20, 2025
DeepSeek came out of nowhere: DeepSeek-R1 — reasoning on par with OpenAI o1, fully open-source under MIT license
June 20, 2024
Anthropic released Claude 3.5 Sonnet — the turning point where Claude began to surpass OpenAI on coding and reasoning (closed source). Claude 3 Opus (March 4, 2024) had set the stage as the first model to match GPT-4 and briefly top the Chatbot Arena
November 30, 2022
OpenAI launched ChatGPT (GPT-3.5) — LLMs went mainstream (closed source)
Jul 31, 2026 · Thinking Machines published A Safe Path to Open Weights — releasing weights indiscriminately isn't safe, but neither is concentrating capable models in a few labs; widen access in stages as the evidence supports.
Jul 27, 2026 · A busy day for the open-weights conversation:
  • Matt Garman, AWS CEO, posted that Amazon has signed the Open Weights and American AI Leadership letter — "Open and closed models are complementary."
  • Dario Amodei, Anthropic CEO, posted Our Position on Open-Weights Models.
  • Moonshot AI released the model weights and technical report of Kimi K3.
Jul 24, 2026 · Jensen Huang joined X and posted his first tweet, sharing NVIDIA's position paper Open Weights and American AI Leadership.
Jul 16, 2026 · Moonshot AI released Kimi K3: Open Frontier Intelligence — 2.8T parameters, 1M context, native multimodal. Russ Salakhutdinov congratulated Kimi founder and CEO Zhilin Yang, his former CMU PhD student: "What a huge win for the open-source community!"
Jul 15, 2026 · Thinking Machines Lab released Inkling, its first open-weights model trained from scratch — a 975B-parameter MoE (41B active) with a 1M-token context window, alongside a preview of Inkling-Small.

Evaluation, Alignment, Interpretability, Trust & Safety

Sep 18, 2026 · The AI Evaluator Forum published "Minimum Conditions for Embedding Evaluators", an open letter signed by 100+ researchers and leaders (including Geoffrey Hinton, Stuart Russell, Arvind Narayanan, Yejin Choi) urging frontier AI companies to embed independent third-party evaluators, with real editorial control, diverse coverage, transparency, retaliation protections, and internal-employee-level access, to assess both AI systems and company safety practices.
Sep 12, 2026 · Anthropic CEO @DarioAmodei published "We Must Pace the Frontier", arguing labs must deliberately slow the growth of AI capabilities — via embedded third-party evaluators, industry-wide safety standards, and international agreements — so that alignment and safety work can keep pace.
Sep 11, 2026 · @janleike: "we may need to give everyone more time for safety and alignment mitigations." Anthropic published its Threat Intelligence Report (AnthropicAI).

Transitioning a probability-based LLM solution into a stable, consistency, production-ready product is not trivial. Last-mile delivery: from innovations to production readiness is important. Achieving production readiness requires extensive work, including rigorous benchmarking, thorough model evaluation, and, most importantly, establishing trust and safety. Some organizations have recognized these challenges and are actively building solutions to close the gap.—examples include Arklex.ai and Virtue AI.

Jun 2026 · @dawnsongtweets: Virtue AI joined Meta Superintelligence Labs (MSL).

Besides, "The Urgency of Interpretability" is clear, opening the black box is necessary.

Jul 2026 · Anthropic Research: A Global Workspace in Language Models.

Academia vs Industry Fusion

First, congratulations to Professors Andrew Barto of UMass Amherst—my graduate alma mater—and his former graduate student Richard Sutton on receiving the 2024 Turing Award this spring for their pioneering work in reinforcement learning.

NeurIPS 2025 in San Diego was a crowded and hugely successful event, showcasing the vibrant intersection of academic research and industrial innovation. The conference highlighted both the opportunities and challenges facing the AI research community. The unexpected incident involving OpenReview has drawn increased attention to fairness, impartiality, transparency, high quality, and privacy protection in academic sharing, prompting concrete actions across the research community, including the NeurIPS donation and the AI Research Leaders: Join Us in Supporting OpenReview initiative.

The ongoing discussion about research vs engineering has intensified this year. We're seeing fascinating position changes between veterans and young researchers, reflecting the evolving landscape of AI careers. 2025 marks a watershed moment as more 95s young researchers are becoming executives at major companies and startups, commanding compensation comparable to NBA stars in Bay Area tech. Notable examples include Shengjia Zhao (Chief Scientist at Meta Superintelligence Labs), Shunyu Yao (姚顺雨) at Tencent, and Fuli Luo (罗福莉) at Xiaomi. These leaders share a common profile: former researchers at OpenAI or top labs like DeepSeek, PhD from Stanford or UC Berkeley, and undergraduate degrees from Tsinghua or Peking University. For startup, we also have new research-type companies like Thinking Machines Lab (founded by former OpenAI CTO, Soumith Chintala, PyTorch co-creator, joined last month) and for corporate, Amazon recently appointing Pieter Abbeel (UC Berkeley professor, co-founder of Covariant) to lead their frontier model research team.

We're witnessing a destructive restructuring of the AI talent landscape. In this new era, everyone needs to find their position. The traditional career paths are being rewritten, and adaptability has become the most valuable skill.

Jul 07, 2026 · Lenny Rachitsky and Noam Segal's 2026 tech worker sentiment survey found the workforce splitting in two: those "amplified" by AI are thriving, while those "destabilized" or "diminished" by it are burned out and pessimistic—now the single strongest predictor of career optimism, ahead of role, seniority, or company size.
Jul 02, 2026 · Zuhayeer Musa, Founder of Levels.fyi, posted "Where (not) to work as SWE in 2026."
June 2026 · The author of "The Cost of Staying" took her own advice, shifting from investor to builder by joining SpaceX / xAI.
May 2026 · Deedy Das, Partner at Menlo Ventures, captured the frenetic SF vibes and the widening divide between the ~10k frontier-AI insiders who hit retirement-level wealth and everyone else navigating layoffs, career paralysis, and an uncertain future of work.
Feb 2026 · An investor at Bloomberg Beta captured a snapshot of this industry transition in her "The Cost of Staying" write-up.

Neo LabsNew · 2026

If 2025 was the year young researchers became executives, 2026 is the year the veterans left to found their own labs. The academia–industry fusion has entered a new phase: instead of joining existing giants, the most senior scientists are spinning up "neo labs" — small, research-first, thesis-driven companies, each betting on a different path beyond today's LLMs.

August 5, 2026
Jeff Dean left Google after 27 years — together with Sanjay Ghemawat and two other senior scientists — to co-found Discovery Loop, a public benefit corporation automating the loop of scientific research: propose, run, evaluate, iterate. Google is a founding investor and cloud partner; in the same reshuffle, Demis Hassabis became chair of Google DeepMind and chief scientist of Alphabet
May 13, 2026
Yuandong Tian (田渊栋), former Meta FAIR research scientist director, emerged from stealth as co-founder of Recursive Superintelligence (RSI) — $650M raised at a $4.65B valuation, betting on self-improving AI systems that autonomously refine their own code and reasoning
November 2025
Yann LeCun left Meta after 12 years as Chief AI Scientist to co-found AMI Labs (Advanced Machine Intelligence) — convinced LLMs are a dead end, betting on JEPA-based world models; raised $1.03B at a $3.5B pre-money valuation in March 2026

Sequoia Capital's Sonya Huang made the same call from the application side. In How Companies Are Building Their Own Intelligence, her opening remarks for Sequoia's "Own Your Intelligence" session, she argued that the endpoint for AI application companies is that every one of them becomes a Neo Lab.


Post-Training RLNew · 2026

Barto and Sutton's Turing Award proved prescient. In 2026, reinforcement learning is no longer just the final alignment step — post-training RL has become the main engine of capability gains. Reinforcement Learning with Verifiable Rewards (RLVR) and GRPO-style recipes, popularized by DeepSeek-R1, are now the standard post-training paradigm across frontier and open-weights labs alike.

The frontier has shifted from static reward datasets to environments: agentic RL trains models inside sandboxes with tools, code execution, and multi-step tasks, where feedback comes from the environment rather than a human label. Environments are becoming the new data — and building good ones is becoming its own discipline (see NVIDIA's agentic RL techniques write-up and the growing systems literature on scaling RL post-training).

The product argument runs alongside the technical one: at that same Sequoia Capital "Own Your Intelligence" session, Fireworks AI CEO Lin Qiao made the case that post-training is how you keep your taste (August 2026).


LLM Native: Open Source Ecosystem in Place

The shift toward GPU-CPU Native infrastructure has accelerated dramatically. Key projects driving this revolution include vLLM, SGLang, llm-d, Nvidia's Dynamo, and LMCache. Looking back at GTC 25, there was no "public" concept of disaggregated inference at all. Just one year later, industry standards have emerged and are growing rapidly. There's also an interesting discussion on customizing popular open-source libraries—check out this article.

Mar 16, 2026 · NVIDIA GTC Keynote 2026, NVIDIA Founder and CEO Jensen Huang shared disaggregated inference on Vera Rubin Prefill + Groq Decode.
Mar 13, 2026 · AWS and Cerebras team up to build new disaggregated inference solution: Trainium Prefill + Cerebras Decode.
2026 · Feb: vLLM started Inferact ($150M seed) · Jan: SGLang started RadixArk (~$400M valuation).

The Cloud Native Computing Foundation (CNCF) turned 10. The Linux Foundation and open-source AI community have remained incredibly active in 2025, driving both agentic AI and foundation model ecosystems forward.

December 2025
MCP 1 Year. Linux Foundation formed Agentic AI Foundation (AAIF) with MCP, goose, and AGENTS.md
October 30, 2025
LMCache joined PyTorch Foundation
October 22, 2025
Ray joined PyTorch Foundation
June 2025
Agent2Agent (A2A) joined Linux Foundation
May 2025
vLLM and DeepSpeed joined PyTorch Foundation
July 2024
vLLM joined Linux Foundation AI & Data

Agentic AI: Moving Fast

The industry consensus is clear: 2025 is the year of the agent. We've seen an explosion of agentic AI frameworks, tools, and applications that are fundamentally changing how we think about AI systems.

From autonomous coding assistants to multi-agent collaboration frameworks, agentic AI has moved from research novelty to production reality. Agent sandbox use cases have also brought Firecracker to the stage. The formation of the Agentic AI Foundation (AAIF) by the Linux Foundation signals that this trend is here to stay.


Expanded Cross-Company Partnerships

Before 2025, top AI labs were tied to one or two companies—OpenAI with Microsoft, Anthropic with Google and Amazon. This year, cross-company partnerships are expanding: OpenAI now works with Microsoft, Oracle, and AWS, while Claude Code is available on OpenRouter. Top AI labs also worked together, in August, OpenAI and Anthropic jointly conducted a pilot Alignment Evaluation Exercise and published research results.


More and More Data Center

Google’s TPU is back in the spotlight, even though it has been around for a long time. Meanwhile, the newly released Trainium 3 and Project Rainier, powered by Trainium 2, have been activated for Anthropic’s Claude model training and inference workloads. And multi-year, billion-dollar agreement between AWS and OpenAI. GPU versus ASIC chips. Undoubtedly, more and more purpose-built physical data centers are being constructed from the ground up, and energy become important.

2026/03/10 · The AI "5-Layer Cake" by Jensen Huang, Nvidia CEO.

Faster and Faster Performance Optimization

The token economy is becoming increasingly evident. More OPTS, shorter TTFT, and longer context lengths, the ranking race in Artificial Analysis continues. Performance and quality (a.k.a. cost efficiency) are becoming critically important at every layer—from applications and agent providers to API and model providers. Companies like Fireworks AI and Together AI did a good job in this domain. AI technologies are innovating fast — inference optimization techniques like speculative decoding are a prime example. Beginning with draft-model speculative decoding (Chen et al., 2023), later the EAGLE (Extrapolation Algorithm for Greater Language‑model Efficiency) series has pushed this approach further: EAGLE-1 (ICML 2024), EAGLE‑2 (EMNLP 2024), EAGLE‑3 (2025), and most recently, SuffixDecoding (NeurIPS 2025 spotlight poster).

Jul 15, 2026 · Fireworks AI announced its Series D: $1.505B raised at a $17.5B valuation, surpassing $1B in annualized revenue and serving 40+ trillion tokens per day.
June 2026 · DeepSeek DeepSpec DSpark.

And all of these new academic innovations are supported by the open-source community from Day 0 for industry adoption. Open-source speed is incredibly fast in today’s AI community.

Mar 2026 · TorchSpec by TorchSpec team, Mooncake team   Nov 2025 · Speculators by vLLM, EAGLE based adaptive speculative decoding by Amazon SageMaker AI team   Jul 2025 · SpecForge by SGLang

Besides, low-level innovations like kernel tuning continue unabated—for example, Perplexity’s October open sourced kernel on GitHub and paper and companies like Cerebras AI design AI-native processors and infrastructure, result in high performance through tight hardware–software combination. Behind high performance is a high-quality dataset.


Pitch Over Performance

Personal experience. Nano Banana did not fully meet expectations, requiring intensive prompting and still leaving room for fine-tuning.


Productivity

Coding: Kiro CLI, Cursor and Claude Code. Doc: ChatGPT. Knowledge: Perplexity.

Jul 16, 2026 · Boris Cherny, creator of Claude Code, mapped the four steps of AI adoption he keeps seeing across teams — one engineer 10x-ing their output with Claude while the rest of the org hasn't caught up: Steps of AI Adoption.
Mar 09, 2026 · Claude Code launches Code Review.
Mar 04, 2026 · The Pragmatic Engineer Podcast: Building Claude Code with Boris Cherny.
Feb 22, 2026 · Claude Code celebrates its 1 year birthday.
Feb 11, 2026 · Cursor Is Dying.

LLM 3 Years

From conversational AI in 2019 to Bedrock in 2023—I become a three-year "veteran" in the LLM domain. OpenAI turns 10 this month: https://openai.com/index/ten-years/. This space is growing at an unimaginable pace, hype and the unknown intertwined. When an industrial revolution is underway, people can only move forward step by step. See @karpathy and @rakyll threads. Recommended interviews: Geoff Hinton & Jeff Dean, Ilya Sutskever, Linus Torvalds, Andrew Ng and Google's blog.


What's Next

No direct answer, but open source will continue to be the way.

Sep 09, 2026 · Two separate Anthropic stories broke the same day:
Jul 28, 2026 · Two visions of the AI future, on the same day:
  • OpenAI voiced support for Pacing the Frontier — a statement from 1,178 employees of frontier AI companies (signatories include Jakub Pachocki, Dario Amodei, Benjamin Mann, John Schulman, Shengjia Zhao, Dawn Song) requesting that the U.S. government support an international effort to develop the technical and governance tools to deliberately pace automated AI development.
  • Mark Zuckerberg, Meta CEO, published the WSJ op-ed "The AI Future Is for Everyone" — superintelligence should broadly empower individuals rather than be centrally controlled by a few institutions.
Jul 14, 2026 · Demis Hassabis, CEO of Google DeepMind: "A Framework for Frontier AI and the Dawning of a New Age." He shared his opinion about AGI: "The magnitude of this technology's impact will be ... 10x of the Industrial Revolution at 10x the speed."
Jul 10, 2026 · Mira Murati, CEO of Thinking Machines Lab: "The Future Worth Building Is Human."
Jul 09, 2026 · AI 2040 — "Plan A," the AI 2027 team's positive scenario for delaying superintelligence until 2040 (announcement).
My own thought The productivity surplus brought by AI could bring today's capital-driven form of society to an early end (capitalism as an economic system emerged in the 16th century, ~500 years ago).
Mar 05, 2026 · Anthropic Economic Research report: Labor market impacts of AI: A new measure and early evidence.
Feb 22, 2026 · A Wall Street article spreading around: The 2028 Global Intelligence Crisis.
Nov 2025 · Project Prometheus founded by Jeff Bezos and Vik Bajaj (ex-Google X), with $6.2B in funding and the slogan "AI for the physical economy."
Apr 03, 2025 · AI 2027 offers a detailed scenario forecast of how transformative AI could unfold over the next few years.