Everything moving in open AI inference, filtered down to your box.
About 150 pull requests a day across the kernel and serving projects, and hundreds of open-weight models a week. Almost none of it changes anything for your box. The titles never tell you which does.
Open weights released this month, with the one thing that decides whether you can use them: whether a GGUF or an NVFP4 exists yet, and whether it fits in your memory at all.
What landed upstream for your silicon — and what quietly left it out, which is the half nobody announces.
People running these models on the same machine as you, with the throughput they measured, how many boxes it took, and the patch that made it work.
Pick your machine
the bar shows what landed — blue worth doing, orange left outThis view is about one machine. Which is yours?
The gap is closing.
Open weights keep closing on the best proprietary models. Those two lines converging are the curve this site is named after.
ProprietaryOpen weights
Late 2024: open weights at roughly 40% of the best proprietary model. Late 2026: within about 15%.
Traced from the Artificial Analysis Intelligence Index v4.3 — their figures, not ours, so the axis carries no numbers. artificialanalysis.ai ↗
The question is whether it reaches your desk.
So every day this reads the merged pull requests, the model releases, the community setups and the papers — and tags each one with the machines it helps, and the ones it leaves behind.
- Every tag shows its evidence — the diff line that justifies it.
- An estimate says so. Memory fit is arithmetic, never a run.
- Unknown stays unknown — never compatible, never false.
So — which machine is yours?
Pick as many as you have. With two or more, every patch says which of your boxes it is for — and which it leaves out.
Now narrow it to your machine
The same week, filtered to what lands on one box — and to what quietly leaves it out, which is the half nobody announces.
Activity across the full collection window
What landed, week by week
Merged pull requests that touch this machine, by the week they landed.
Where the work lands
Which part of the stack is moving for this machine.
Find your next model
More filters
Explore the comparison chart Size, speed, context & downloads
Size against measured speed
Only models somebody actually ran appear here: the vertical axis is a measured figure from a community recipe, not an estimate.
Does your setup need a change?
Corrections address a reported problem; improvements are optional. Nothing here knows which fixes your installed version already has.
1. Match your setup
2. Decide what applies
Compatibility exclusions
Changes that leave this machine outCommunity recipes
Real setups by hardware, engine and quantisation, each with the author's own conditions.
Find a setup
Where this is going
How it got here: which machines are gaining people who publish for them, how fast the numbers move, and how much of what ships fits a box you can buy.
What you get on this machine
What to expect, against what gets published
Every attributed single-box figure, per machine: the bar is the middle half, the thick line the median, the ring the record. The record is not the expectation — on a DGX Spark it is five times the median. Multi-box figures are in the tooltip, not the drawing.
What people run it with
Engines named by published setups, per machine. Not one preference: a DGX Spark is mostly vLLM, a Strix Halo almost entirely llama.cpp.
Which knobs actually get turned
Techniques named by setups for your machine — what people did, not what papers propose.
Over time
Who is publishing setups, and for which machine
One bar per month, split by the machine the setup covers; one covering three machines counts in all three. The date is the repository’s, not the figure’s — when a number was published is recorded nowhere.
Machine by machine, month by month
The same setups, one row per machine instead of one pile. Each row is scaled to its own busiest month, printed on the right, so every shape is legible and heights across rows mean nothing. Totals do not compare either: each machine is found by its own search terms, so a row compares its present against its own past.
When each machine got its first published setup
Ordered by when the first setup appeared; the number on the right is the total. Machines nobody has published for are counted underneath, not listed.
Published figures over time, and the ceiling they kept moving
Every attributed single-stream figure, and the running best — the record as it stood that day, not today’s. Aggregates, prefill and training figures are excluded, by the same rule the records page uses.
Merged patches, week by week
Only changes worth doing, by the part of the stack they touch.
Papers worth reading, day by day
arXiv submissions that cleared the bar, by area.
What the daily record has collected so far
The catalogue and the usage figures arrive without history. Every pass appends a row, so this one starts empty and fills itself from here.
Who publishes this
Who to follow for your machine
Two counted quantities, no ranking: when they last published, and how many figures they measured themselves. With a machine picked, the accounts publishing for it are in colour. Up and to the right is active and measuring.
Whose numbers are their own
For the busiest accounts: figures they measured, figures quoted from elsewhere, and setups reporting none. A setup without a figure is not worse, it is a different thing.
New faces against the regulars
Accounts publishing each month, split by whether it was their first. A growing community and the same twenty people publishing more give the same total curve; only this separates them.
The shape right now
The machines: memory against bandwidth
The two numbers that decide what you can load and how fast it decodes, on log scales. Up and to the right is fast and roomy; bottom right is the memory-rich, bandwidth-poor corner. Every dot answers on hover.
The catalogue by size, against what each box can hold
Dashed lines are memory ceilings at 4-bit, weights × bytes × 1.2. Arithmetic, not a measurement: it knows nothing about KV cache, fragmentation or whether a kernel exists. Only your machine’s ceiling is labelled; the rest answer on hover.
Weight formats
What exists, and how much of it your machine could load at all.
How much of what fits has anyone actually run
The bar is models with a published figure. “Fits” is not the bar: that is arithmetic on the weights, printed as a number. The faint tail is what was measured for the first time in the last 90 days.
Families by releases
Checkpoint count. This one measures who republishes: requantisations, repacks and derivatives all count as releases.
Families by downloads
The same families, ordered by use. High on the left and low here is a repackaging shop, not a lab.
Papers by area, and what they release
“Does not say” is the majority, and it is its own state: a paper silent about its code has not told you there is none.
Techniques named in papers
Mentions inside a 30-day window: what got written about, not what got adopted.
The negative control: a plain grep against the tagger
How many patches a keyword search would have marked for each machine, how many the tagger marked, and where the two disagree.
Papers
Narrow the reading list
Can I run it?
Narrow the matrix
Device count, engine, quantisation and context differ between runs. Figures exclude prefill and concurrent streams. A missing run is missing evidence, not incompatibility.
Find the right model for the task
Match a task to your hardware
Each score belongs to the model and setup it was measured in; a community recipe may use another quantisation or context. Memory fit is an estimate. No published run is missing evidence, not incompatibility.
Index
Apps
Usage now, and where it is going
Area is tokens routed in the last 30 days; colour is those 7 days against the three weeks before. Both published by OpenRouter, never combined.
Narrow the ranking
First seen this week
Creators
Find a creator
What the signal is, and what it is not
Following
Compare candidates
Every cell says what produced it. Scores from different evaluations are never averaged into one number, a missing value is shown as missing and not as zero, and a run on more devices than you have is marked.
What is being compared
What’s new for you
Have an agent watch this for you
The same updates answer over MCP — filtered by your setup, each one carrying the reason it is here — so you can ask what changed for my machines this week from wherever you already work, instead of coming to look.
Personalize your feed
Put an agent on this setup
The catalogue answers over MCP as well as on screen, so an agent can ask what this page shows: what changed for your machines, whether a model fits one of them, whether anyone just beat the best published figure for your box. With a token it reads the setup above by itself — you never describe your hardware twice — and it can queue something for the daily pass without leaving the conversation.
Saved
Things to try and keep for later.