AI is built in two layers: a handful of frontier labs whose models you rent, and an open-weights layer, a few months behind, whose models anyone can download and run.
- The frontier list is short because of cost. Training at the leading edge takes capital, power and physical infrastructure on the scale of heavy industry, not secret know-how.
- Open weights trail by months, not years. What the frontier reaches soon becomes downloadable, which works against lasting concentration of power and means safety work at the frontier cannot carry the whole load.
- This page keeps no leaderboard on purpose. Rankings go stale within weeks, rest on benchmarks researchers have shown can be gamed, and say little about whether a tool works on your task.
- Watch consolidation and the gap, not the race. One frontier lab has already been folded into a much larger company, and whether the open layer catches up or falls behind decides how concentrated AI becomes.
Two layers. At the frontier, a handful of labs with the capital to train the largest models. Beneath them, an open-weights wave publishing models anyone can download and run — trailing by months, not years. Almost every argument about AI power concentration is really an argument about the gap between those two layers, and whether it is widening or closing.
The frontier tier
This page groups five labs as the frontier tier. Four of them ship models at the leading edge of general capability; the fifth, Meta, is on the list with the weakest claim, because the closed models it has actually shipped are described as not frontier-level (see its entry). An earlier version of this page said the tier had widened from three labs to five in mid-2026, citing unnamed trackers; no source for that count could be found, so it is withdrawn.Reasoning — this page’s own grouping, not a count published by any tracker, amended 17 Sep 2026. No tracker read here defines a frontier tier by lab. Names and standings change quickly; corrections are made in place on this page.
The lab that made this a public phenomenon rather than a research topic. Its November 2022 launch is the starting line for nearly every measurement on this site.
Founded by former OpenAI researchers; positions its work around alignment and published behavioural principles. Also, in the interest of disclosure: the tooling behind parts of this site.
The deepest research lineage of the five — the transformer paper itself came out of Google in 2017Vaswani et al., "Attention Is All You Need", submitted 12 June 2017 — arXiv:1706.03762, read at source 11 Sep 2026. That abstract page lists the eight authors but carries no affiliation line, so the Google attribution is not verified on it. — plus the rare advantage of owning its own chips and datacentres.
Spent years as the largest publisher of open-weights models, then paused new open-weight Llama releases and moved its flagship to a closed line — Muse Spark, of which Meta says only that it hopes “to open-source future versions of the model”. It has not left open weights entirely: on 10 August 2026 it released Muse Glimmer, “a 30-billion-parameter model”, with weights under “a permissive Apache 2.0 license”. The shift is still worth watching, because the model at the top of Meta’s line is the closed one. Its place on this list is the weakest of the five: reporting read in September 2026 describes the closed models Meta has actually shipped as not frontier-level.Read at source 11 Sep 2026. Spyglass, 10 Aug 2026 — Meta is "setting aside that notion of ‘open’ in favor of a closed approach to AI development", and its Muse Spark 1.2 is "by most accounts a very good, but not frontier-level model". Digital Applied, H1 2026 retrospective — "Meta ship-paused open-weight Llama after pivoting frontier attention toward the closed Muse line." Amended 11 Sep 2026: this entry previously said Meta "moved toward paid frontier offerings"; no source read here places a shipped Meta model at the frontier. Read at source 16 Sep 2026: Meta, Introducing Muse Spark, 8 Apr 2026; Meta, Introducing Muse Glimmer, 10 Aug 2026. Amended 16 Sep 2026: this entry previously said only that Meta “pivoted to a closed line”, which left out an open-weights release made the same day as the Spyglass report it cited.
The newest of the five, and the one whose corporate structure changed most in 2026: SpaceX announced its acquisition of xAI on 2 February 2026 in an all-stock deal, reported as the largest corporate merger to date at a combined valuation of about $1.25 trillion.Read at source 11 Sep 2026. Deal date and structure: Value Add VC — "On February 2, 2026, SpaceX acquired xAI in an all-stock transaction that valued the combined company at $1.25 trillion". Valuation split confirmed by Motley Fool, 31 Mar 2026: "The merger valued SpaceX at around $1 trillion and xAI at $250 billion, giving them a combined value of $1.25 trillion." The "largest merger to date" framing rests on CNBC’s headline, "Musk’s xAI, SpaceX combo is the biggest merger of all time, valued at $1.25 trillion" — not read at source, cnbc.com returns 403 to automated requests. A live corporate situation — expect this entry to date faster than the rest of the page.
The open-weights wave
One stratum below the frontier sits the layer that keeps this from being a five-company story. These labs publish trained weights so anyone can download, run and modify the model on their own hardware — no API and no per-token bill, though each licence sets its own conditions.Read at source 11 Sep 2026: Digital Applied, open-weight models H1 2026 — "Qwen ran the most active release cadence (3.5 family in February, 3.6 in April)"; "DeepSeek shipped a single architectural reset with V4 Preview on April 24". That retrospective covers DeepSeek, Qwen and Llama only — it does not cover Moonshot (Kimi) or Mistral. Both, with their licences, are checked at source on AI labs outside the US (18 Sep 2026).
The strategic fact worth carrying: the gap is measured in months, not years. Capability that costs hundreds of millions to reach at the frontier becomes downloadable within a year or so. That is simultaneously the strongest argument against permanent concentration of power, and the strongest argument that safety measures at the frontier cannot be the only line of defence — because the capability does not stay there.
Do not judge who controls AI by the frontier labs alone. What they build soon reaches a layer anyone can download, so both the worry about concentration and the work on safety have to account for it.
Why the list is short
Not secrecy. Cost. A frontier run occupies tens of thousands of accelerators continuously for months; add the power draw, the cooling, the hardware itself, and the bill runs into the hundreds of millions before a single user sees the result. The International Energy Agency projects datacentre electricity demand more than doubling by 2030 to around 945 TWh, with AI the most important driver of that growth.IEA, World Energy Outlook Special Report: Energy and AI, April 2025 — read at source 11 Sep 2026 in the report PDF (iea.org itself refuses automated requests): "Data centre electricity consumption is set to more than double to around 945 TWh by 2030. This is slightly more than Japan’s total electricity consumption today. AI is the most important driver of this growth, alongside growing demand for other digital services." Corrected 11 Sep 2026 from "roughly doubling" and "fastest-growing component". See the six layers for the physical substrate.
Which means the real constraints on who builds AI are the same constraints that govern heavy industry: capital, power, and physical infrastructure. That is an unglamorous answer, and it is the correct one.
The release pace
One public release tracker counts 249 frontier models between ChatGPT's launch and September 2026 — a count that depends entirely on where each tracker draws the line around “frontier” — which works out at more than one significant release per week across the industry. The practical consequence for a reader: any page ranking models against each other is stale almost immediately, which is why this site records who is building and at what cost rather than maintaining a leaderboard.AI Release Tracker, read 17 Sep 2026: “We cover 249 tracked frontier models from Anthropic, OpenAI, Google, Meta, SpaceXAI, DeepSeek, Mistral, Moonshot AI, Z.ai, Qwen, NVIDIA, starting with the launch of ChatGPT on November 30 2022.” 249 over the roughly 198 weeks since then is about 1.3 a week. Another tracker draws the line far tighter: llm-releases.com, read 11 Sep 2026, counted “349 total models” of which “75 Frontier”. This page previously gave 198 for July 2026 with no tracker named; that figure could not be traced and is withdrawn.
Why this page has no leaderboard
Model rankings are the most-searched thing in this subject and the least durable. This site does not keep one, and the reasons are worth stating rather than implying.
They expire. With roughly one significant release a week across the industry, any ranking is stale before most readers find it. A page that is wrong by the time it is read is not a reference; it is decoration.
The measurements underneath them are contested. Benchmarks are scored right-or-wrong with no credit for admitting uncertainty, which rewards confident guessing over honesty — the mechanism behind why models make things up. Test material circulates publicly, so scores can rise without ability rising. And in April 2026 researchers showed that all eight of the leading agent benchmarks they audited could be driven to near-perfect results without solving a single task, by attacking the evaluation harness rather than the problem.Berkeley RDI, "How We Broke Top AI Agent Benchmarks", April 2026 — read at source 11 Sep 2026: rdi.berkeley.edu, "every single one" of the eight "can be exploited to achieve near-perfect scores without solving a single task"; "Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores." Scores reached: 100% on Terminal-Bench, SWE-bench Verified and Pro, WebArena, FieldWorkArena and CAR-bench; ~98% on GAIA; 73% on OSWorld. Narrowed 11 Sep 2026 from "every major agent benchmark" to the eight actually audited. Detail on how AI is tested and AI agents.
Rank does not predict usefulness. The gap between a headline score and behaviour you could depend on has been measured at tens of points: in one enterprise-agent study, agents scoring 60% on a single run scored 25% once the same task had to succeed eight times running."Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems" (CLEAR), arXiv:2511.14136, 2025 — read at source 11 Sep 2026: "inadequate reliability assessment where agent performance drops from 60% (single run) to 25% (8-run consistency)". Made specific 11 Sep 2026; "tens of points" previously stood with no source named. What decides whether a tool works for you is the task, the context you supply and the checking you do — none of which appear in a ranking.
And the ordering is unstable at the top. Where leading models differ by a point or two on a contested measure, the ranking is noise presented as a finding.
Ask which layer you are choosing between — a frontier model you rent, or an open-weights model you hold. That choice is durable, consequential and unlikely to be reversed by next week's release. Everything else is a preference you can test in ten minutes on your own actual work, which is a better evaluation than any leaderboard because it uses your data and your standards. The choice itself is weighed in open or closed.
Skip the rankings. Decide first whether you want a model you rent or one you hold, then try the candidates on your own work, which tells you more than a leaderboard can.
The competition, described carefully
The rivalry between labs gets covered as a war, and the metaphor does real damage — it implies a finish line, a winner, and a scoreboard, when the observable dynamics are duller and more consequential.
What is actually happening is consolidation into larger parents. One of the five frontier labs ceased to exist as an independent company in 2026, absorbed into a much larger industrial group. Another, once the largest publisher of open-weights models, moved its flagship to a closed line while still releasing some open weights. Both shifts move capability and accountability, and neither is a battle.
The binding constraint is physical, not strategic. Chips, electricity, cooling and capital decide who can train at the frontier — which is why the list is short and why the energy question is a competition question. Nobody out-thinks a power grid.
And the gap that matters is not between labs. It is between the frontier and the open layer beneath it, currently months wide. Whether that gap closes or widens decides how concentrated this technology becomes — a far more important number than which company is briefly ahead. Epoch AI tracks it as a trend.Epoch AI, Open models lag state-of-the-art closed models by 4 months, read at source 17 Sep 2026: “Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in the Epoch Capabilities Index (ECI)”.
WHAT THIS PAGE DELIBERATELY DOES NOT DO
No rankings, no "best model" verdict, no benchmark tables, and no framing of the industry as a race with a winner. Those are obsolete within weeks and they are the most-copied content on the internet. What holds longer is the structure: two layers, a months-wide gap, a cost floor that decides who can play, and a consolidation trend that decides who answers for the result.
Concentration of this much capability in a handful of organisations is worth taking seriously — and the open-weights layer is a genuine, under-reported counterweight to it. Both things are true. The honest position is neither "a few companies own the future" nor "it is all open and fine": it is that the gap between those layers is the number to watch, and it is currently measured in months.
When you read AI industry news, look past who is briefly ahead. The two things that change who holds this technology are consolidation into larger parents and whether the open layer is closing on the frontier.