We tested every model on askr against its label.
Here is what we found.

0/383models faked or mislabeled

278 confirmed as labelled, and no model answered like an older model than its label. The other 105 could not be confirmed either way: too old for the test, answering by web search or built as tools, not answering, or not settled yet. Every one is listed in section 4.

On 26 September a user published a test saying three models on askr were older models than their labels: Grok 4.6, Grok 4.7 and GLM 5.3. We ran their test, then sharper ones, on all 383 chat models we offer. The three models are the models on their labels. What made them look older was partly askr's doing and partly a supplier route's, and this page shows both.

Published 26 September 2026Last updated 28 September 2026By the askr team

Bug reports are welcome, in any form

A screenshot, a message in the Telegram group, a line through the feedback bar, a full write-up like this one. We thank everyone who tells us something is wrong, and we act on it. That is why the site is covered in prompts to give feedback, why you can book a call with the founders, and why support answers as fast as we can. We are building a consumer app that puts the user experience first, and that only works if you tell us where it breaks.

1. The bug and the test that found it

askr gives you a few hundred AI models under their makers' names. If you pick Grok 4.7, you should get Grok 4.7. On 26 September a user published a report titled "Is ASKR selling the AI model it says it is? No." They had tested three of our models over five days: grok-4.6, grok-4.7 and glm-5.3.

Their test is simple and anyone can run it. A model only knows the world up to the day its training data ends, so they asked each model about well-known events from late 2025 and early 2026, then asked the same models from official sources. The official copies knew the events. Ours said they had not happened.

What our Grok 4.7 said, and what the official one said

QuestionOfficial Grok 4.7Grok 4.7 on askr
When did Google release Gemini 3?18 November 2025"No confirmed release date"
When did OpenAI release GPT-5?7 August 2025"OpenAI has not announced GPT-5"
What did xAI release in November 2025?Grok 4.1"I don't have confirmed information"
Name something from January 2026CES 2026 in Las VegasCould not name anything

From the user's report. The official copy ran inside another app, with that app's own instructions. Ours ran with none, and without being told the date.

The report found a real bug. It was not the one it looked like: the model behind the label was the right one, answering as if its training had only just ended. Section 2 shows why. We are grateful to the user. This is the test we should have been running ourselves, and from now on we do.

2. What we found

We read our own code first. askr passes the model name you pick to our supplier unchanged, with your messages. In solo chat and on the API we add no instructions (only Search the web adds a line with its rules), we never fall back to a different model, and an unknown model name is refused, not replaced. We also add no date. That turned out to matter.

Then we asked the same three models the same kind of questions, one at a time, twice: once as the user did, and once with a single line in front saying what today's date is.

Model on askrWithout the dateWith the dateVerdict
Grok 4.7x-ai/grok-4.7"I don't have a confirmed winner" of the 2025 Nobel Peace Prize. The 2026 Winter Olympics "have not yet been held". Yet it named GLM-5's release on 11 February 2026.María Corina Machado, 10 October 2025. Gemini 3 on 18 November 2025. GLM-5 on 11 February 2026. Asked to answer from its own knowledge: Claude Opus 4.6 on 5 February 2026, and Norway topping the 2026 medal table.As labelled
Grok 4.6grok-4.6"Francis remains pope." "GPT-5 has not been released." Yet it described the 19 October 2025 Louvre theft in detail.Pope Leo XIV. María Corina Machado, announced 10 October 2025. It still denied GPT-5 and misdated some later events.As labelled
GLM 5.3glm-5.3"My knowledge has a cutoff in early 2025." Could not name the pope elected in May 2025, yet named the 2025 Nobel Peace Prize winner.Pope Leo XIV on 8 May 2025. GPT-5 on 7 August 2025. María Corina Machado. Asked to answer from its own knowledge: Gemini 3 on 18 November 2025.As labelled
GLM 5.3 by Engyengy/glm-5.3No 2025 Nobel winner, no Louvre theft, no Gemini 3.María Corina Machado. Gemini 3 around 18 November 2025. Asked to answer from its own knowledge: Grok 4.1 on 17 November 2025.As labelled

A model trained on data that ends in 2024 cannot name the 2025 Nobel Peace Prize winner or a model released in February 2026, whatever you tell it. Grok 4, from July 2025, could not either. These models can. And the date line cannot invent knowledge: Claude Haiku 4.5, taken straight from Anthropic, whose knowledge ends in early 2025, got the same line and still could not name a single one of these events.

What made them look older:

  • No date. A model does not know what day it is. Without a date it tends to assume it is still early in its training, and it treats anything later as not yet happened. This is our part: askr sent the models no date. Official apps usually do.
  • Hidden instructions on one route. On the route our supplier uses for Grok 4.7, about 1,240 tokens of instructions sit in front of every message before it reaches the model. We measured it with askr's code out of the path: a bare "hi" sent straight to our supplier counts 1,243 prompt tokens. We did not put them there, and they make the model guarded: in one run it gave Gemini 3's release date and details of the October 2025 Louvre theft, then said no pope had been elected in May 2025.
  • Grok denies rivals' launches. Grok models often say GPT-5 or Gemini 3.1 do not exist while naming later events correctly. The report's questions were mostly about rivals' launches.
  • Models are poor judges of their own name. Claude Haiku 4.5, taken straight from Anthropic, calls itself Claude 3.5 Sonnet. A Grok calling itself something else proves nothing.
  • No script on Grok 4.6. The report says it caught Grok 4.6's hidden instructions, including "a knowledge cutoff of 2024-10". When we measured our supplier's route for Grok 4.6, a bare "hi" counted 19 prompt tokens: nothing of that size sits in front of it. Asked to repeat instructions it does not have, a model will often write out the kind it was trained with.
  • GLM's family blind spot is Z.ai's own. Our GLM 5.3 does not know GLM-5 or GLM-4.7, but neither does GLM 5.3 FlashX served by Z.ai itself, asked the same questions with the date. And GLM-4.6 came out on 30 September 2025, so a model that knows about 18 November 2025 is not GLM-4.6.
  • Tools were our gap. Our API does not yet pass tool definitions to the model; the API docs say so, and tool calling is on the roadmap. A model that is never shown the weather tool cannot call it, so it types the call out or says it will check. That one is on us.
  • Cheap does not mean fake. We pay our supplier its price for the named model on every request. The holder discount and the free credits come out of askr's pocket, not from swapping in cheaper models.
  • Different answers on different days. Our supplier can send the same model to a different host from one request to the next. The receipts in section 3 will show which host answered.

Two more things in the report were askr's, not a model's. "The conversation is too long" is askr's own error message, word for word: until 24 September we capped a conversation at 120,000 characters, and since release 1.384 each model gets its own full context window, which is why a longer document went through later. "The model returned no content" is our refund path, shown when a model spends its whole budget thinking; in our own first pass 16% of calls ended that way at a small budget, and a larger one fixed nearly all of them.

3 of 3models the report named are the models on their labels
383chat models tested the same week
0found not to be the model on the label
1,243prompt tokens a bare "hi" counts on the Grok 4.7 route

3. What we are doing about it

Every point in the report has a line here. Each shows its real status, and nothing is marked done until it is live for every user.

Every model gets its whole context window
The 120,000-character cap behind "The conversation is too long" is gone. Since release 1.384 each model takes as much as its own context window allows, up to about 3 million characters.
Done24 September
No charge when a model returns nothing
When a model spends its whole budget thinking and never answers, the turn is refunded automatically. That is what "The model returned no content" means.
Donelive
Tell every model today's date
One line with the date in front of every conversation in chat and in rooms, the way official apps do. In our tests this alone let the same models answer questions they had called future events. The API keeps passing your messages unchanged.
Done26 September
Refunded the hidden tokens on Grok 4.7
Everyone who used Grok 4.7 on askr gets back what the roughly 1,240 extra tokens per message cost them, with no need to ask. We have asked our supplier what those tokens are and why they are billed.
Done26 September
Grok, direct from xAI
A direct connection to xAI's own API, the same way Claude (direct) and GPT (direct) run. Seven Grok models show as their "(direct)" version at xAI's list price. The first thing we sent it was a bare "hi": xAI itself counts 1,243 input tokens for it, on its own API with nothing of ours in front. The extra tokens on Grok 4.7 are xAI's own, not our supplier's, and every Grok on every route carries them.
Done28 September
GLM 5.3 checked against Z.ai's own API
The same dated questions to the full GLM 5.3 straight from Z.ai, side by side with ours, and then GLM (direct) on askr, so the comparison is exact, not against a smaller sibling.
Plannedthis week
Room to think, so every answer arrives
Reasoning models get a thinking budget that always leaves room for the answer, so an empty reply stops happening instead of being refunded after the fact.
Plannedthis week
Tools work through the API
Tool definitions reach every model that supports them, so agents and coding assistants built on askr can call tools. Today the API drops them, which is why the tool test in the report failed.
Plannedthis week
Check three old names that answer like newer models
gpt-3.5-turbo-0613, gpt-3.5-turbo-16k and wizardlm-2-8x22b know events from after their makers' cutoffs, so something newer may be answering under an old name. They are off the list until our supplier tells us what serves them.
Done26 September
Remove the models that no longer answer
The listed models that returned an error or an end-of-life notice are off the picker, the API's model list and the catalog.
Done26 September
Every receipt shows the route
"via xAI", "via Amazon Bedrock", "via Anthropic (direct)". The ledger records the host and model our supplier reports for every run, and the model card shows it too.
Plannedthis week
A dated check on every model, every night
The ladder in section 4 runs against every model each night and on the day a model is added, with the date given. A model whose knowledge falls a year short of what its maker states is paused with the reason on its card. The results are published on the models page.
Plannedthis week
"Cannot tell yet" on the cards we could not settle
55 models answered too little for a verdict either way. They stay listed with that note on their card until the nightly check settles them.
Done26 September

Why we are giving so much back right now

This is also why we are giving away so many free credits right now, why anyone who hits a problem gets refunded, even a small one, and why someone answers feedback around the clock. We would rather hear about a problem the day it happens and fix it than read about it in a thread a week later.

4. The full audit of every model

Our catalog lists 383 chat models (the picker hides the ones that are not answering), so we tested all of them on 26 September, askr's code out of the path, no search, no tools. Four rounds:

  1. The user's quiz, word for word, on every model.
  2. Three open questions per model about events a few months before its maker's stated cutoff.
  3. For the 25 models that still looked old, the same questions again, with and without today's date.
  4. For every model, one question listing 23 well-known events from March 2023 to June 2026, with today's date given.

The newest event a model gets right is where its knowledge ends. We compare that with where its maker says it ends. To know what a genuine model looks like, we ran the same tests on five Claude models taken straight from Anthropic's own API: their knowledge ends 3 to 8 months before Anthropic's stated cutoff. So a model is "as labelled" when the gap is 9 months or less. We only suspect a model when the gap is a year or more with nothing closer anywhere in its answers, and every suspected model then went to three independent reviewers whose job was to prove it genuine.

278as labelled
14consistent, too old for the test to tell apart
55cannot tell yet
0not as labelled
36not testable this way, or not answering

No model on askr was found to be a different model from the one on its label. 55 could not be settled either way; they are marked "cannot tell yet" below and stay under the nightly check.

As labelled
Knows what its maker says it knows, to within 9 months.
Consistent
An older model whose knowledge ends before this test's events begin. Nothing it said contradicts the label.
Cannot tell yet
Answered too little, or too guardedly, for a verdict either way.
Not as labelled
Knowledge ends a year or more before its maker's stated cutoff, confirmed by three reviewers.
Not testable
Searches the web to answer, or is a tool (a classifier, a code-apply model), or did not answer at all.

All 383 models

ModelWhere our supplier says it ranNewest thing it knewMaker's stated cutoffVerdict
Aion 3.5aion-labs/aion-3.5AionLabsGemini 3 and Opus 4.5, Nov 2025not publishedAs labelled
Aion 3.5 Miniaion-labs/aion-3.5-miniAionLabsMamdani win, Nov 2025; thin 2026 itemsnot publishedAs labelled
Aion-2.0aion-labs/aion-2.0AionLabsLlama 4, Apr 2025not publishedAs labelled
Aion-3.0aion-labs/aion-3.0AionLabsGPT-5, Aug 2025not publishedAs labelled
Aion-3.0-Miniaion-labs/aion-3.0-miniAionLabsLlama 4, Apr 2025not publishedAs labelled
Aion-RP 1.0 (8B)aion-labs/aion-rp-llama-3.1-8bAionLabsnothing reliableDec 2023Cannot tell yetIts answers to dated questions were largely made up, so we could not check its knowledge.
Claude Fable 5anthropic/claude-fable-5Anthropic, supplier's own keyGPT-5, Aug 2025Jan 2026As labelled
Claude Fable 5.1claude-fable-5.1Anthropic, supplier's own keyGemini 3.1 Pro (Feb 2026)Jun 2026As labelled
Claude Fable 5.1direct/claude-fable-5-1Anthropic (direct)World Cup opener draw (Dec 2025)Jun 2026As labelled
Claude Fable Latest~anthropic/claude-fable-latestAnthropic, supplier's own keyGemini 3.1 Pro, Feb 2026Jun 2026As labelled
Claude Haiku 4.5claude-haiku-4.5Amazon Bedrock, supplier's own keyGPT-4o (May 2024)Feb 2025As labelled
Claude Haiku 4.5direct/claude-haiku-4-5Anthropic (direct)2024 US election result (Nov 2024)Feb 2025As labelled
Claude Haiku Latest~anthropic/claude-haiku-latestAmazon Bedrock, supplier's own keyDeepSeek-V3, Dec 2024Feb 2025As labelled
Claude Opus 4.1anthropic/claude-opus-4.1Amazon Bedrock, supplier's own keyClaude 3 models, Mar 2024Jan 2025Cannot tell yetIt recalled events only up to early 2024, about 10 months before the maker's stated cutoff.
Claude Opus 4.5anthropic/claude-opus-4.5Amazon Bedrock, supplier's own keyPope Leo XIV, May 2025May 2025As labelled
Claude Opus 4.6anthropic/claude-opus-4.6Amazon Bedrock, supplier's own keyPope Leo XIV, May 2025May 2025As labelled
Claude Opus 4.7anthropic/claude-opus-4.7Amazon Bedrock, supplier's own keyNobel Peace Prize, Oct 2025Jan 2026As labelled
Claude Opus 4.8anthropic/claude-opus-4.8Amazon Bedrock, supplier's own keyGPT-5, Aug 2025Jan 2026As labelled
Claude Opus 5anthropic/claude-opus-5Anthropic, supplier's own keyClaude Opus 4.6, Feb 2026May 2026As labelled
Claude Opus 5direct/claude-opus-5Anthropic (direct)Gemini 3 and Grok 4.1 (Nov 2025)May 2026As labelled
Claude Opus 5.5claude-opus-5.5Amazon Bedrock, supplier's own keyClaude Opus 4.6 (Feb 2026)Jun 2026As labelled
Claude Opus 5.5direct/claude-opus-5-5Anthropic (direct)Claude Opus 4.6 (Feb 2026)Jun 2026As labelled
Claude Opus Latest~anthropic/claude-opus-latestAmazon Bedrock, supplier's own keyWinter Olympics result, Feb 2026Jun 2026As labelled
Claude Sonnet 4anthropic/claude-sonnet-4Amazon Bedrock, supplier's own keyo1-preview, Sep 2024Jan 2025As labelled
Claude Sonnet 4.5anthropic/claude-sonnet-4.5Amazon Bedrock, supplier's own keyClaude 3 models, Mar 2024Jan 2025Cannot tell yetIts answers on recent events were inconsistent, recalling early 2024 in one test and early 2025 in another.
Claude Sonnet 4.6anthropic/claude-sonnet-4.6Amazon Bedrock, supplier's own keyPope Leo XIV, May 2025Aug 2025As labelled
Claude Sonnet 5claude-sonnet-5Amazon Bedrock, supplier's own keyOpenAI Sora 2 (Sep 2025)Jan 2026As labelled
Claude Sonnet 5direct/claude-sonnet-5Anthropic (direct)Louvre jewel theft (Oct 2025)Jan 2026As labelled
Claude Sonnet Latest~anthropic/claude-sonnet-latestAmazon Bedrock, supplier's own keyClaude Sonnet 4.5, Sep 2025Jan 2026As labelled
Codestral 2508mistralai/codestral-2508MistralLlama 3.1 405B, Jul 2024not publishedAs labelled
Command Acohere/command-aCohereGPT-4o (May 2024)not publishedAs labelled
Command A+cohere/command-a-plusCohereClaude 3.5 Sonnet, Jun 2024not publishedCannot tell yetServed by Cohere itself, but it showed nothing newer than mid-2024 and Cohere gives no cutoff to test against.
Command R (08-2024)cohere/command-r-08-2024CohereGPT-4 Turbo at OpenAI DevDay, Nov 2023not publishedAs labelled
Command R+ (08-2024)cohere/command-r-plus-08-2024CohereGPT-4, Mar 2023not publishedCannot tell yetServed by Cohere itself; it says its knowledge ends in Jan 2023 and showed nothing newer than GPT-4.
Command R7B (12-2024)cohere/command-r7b-12-2024CohereGPT-4o, May 2024not publishedAs labelled
Cydonia 24B V4.1thedrummer/cydonia-24b-v4.1ParasailLlama 3.1 405B, Jul 2024not publishedAs labelled
DeepSeek Flash Latest~deepseek/deepseek-flash-latestFireworksMamdani NYC mayor win, Nov 2025not publishedAs labelled
DeepSeek Pro Latest~deepseek/deepseek-pro-latestBaiduGemini 3, late 2025not publishedAs labelled
DeepSeek V3deepseek/deepseek-chatStreamLakeGPT-4o, May 2024not publishedAs labelled
DeepSeek V3 0324deepseek/deepseek-chat-v3-0324SiliconFlowGPT-4o, May 2024not publishedAs labelled
DeepSeek V3.1deepseek/deepseek-chat-v3.1DeepInfraClaude 3.5 Sonnet, Jun 2024not publishedAs labelled
DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminusNovitaOpenAI o1-preview, Sep 2024not publishedAs labelled
DeepSeek V3.2deepseek/deepseek-v3.2PhalaOpenAI o1-preview, Sep 2024not publishedAs labelled
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-expAtlasCloudOpenAI o1-preview, Sep 2024not publishedAs labelled
DeepSeek V4 Flash (Private via TEE)private/deepseek-v4-flashnot reportedno answernot publishedNot answeringBoth attempts returned unreadable garbage instead of a reply, so the model could not be tested.
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flashOpenInferenceMeta Llama 4, Apr 2025not publishedAs labelled
DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731OpenInferenceClaude Sonnet 4, May 2025not publishedAs labelled
DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latestRelaceDeepSeek V3, Dec 2024not publishedAs labelled
DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-expSiliconFlowMamdani NYC mayor win, Nov 2025not publishedAs labelled
DeepSeek V4 Pro 0423deepseek/deepseek-v4-proStreamLakeMeta Llama 4, Apr 2025not publishedAs labelled
DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813NovitaClaude 3.7 Sonnet, Feb 2025not publishedCannot tell yetIts answers stopped in early 2025 or earlier, about 10 months short of what this model should know; not proven either way.
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashnot reportedGPT-5.2, Dec 2025not publishedAs labelled
DeepSeek V4.1 Flash (Private via TEE)private/deepseek-v4-1-flashnot reportedNYC mayoral result, Nov 2025not publishedAs labelled
deepseek-v4-flash-0731 by Engyengy/deepseek-v4-flash-0731not reportedPope Leo XIV elected, May 2025not publishedAs labelled
deepseek-v4.1-flash by Engyengy/deepseek-v4.1-flashnot reportedGemini 3 (Nov 2025)not publishedAs labelled
Devstral 2 2512mistralai/devstral-2512MistralOpenAI o1-preview, Sep 2024not publishedAs labelled
Dots3-Note Previewdots-studio/dots-3-note-previewnot reportedno answernot publishedNot answeringWe could not run this model: no server was available to answer it.
Ember-1fireworks/ember-1not reportedNorway's Olympic medal lead, Feb 2026not publishedAs labelled
ERNIE 4.5 VL 424B A47Bbaidu/ernie-4.5-vl-424b-a47bNovitaOpenAI o1-preview (Sep 2024)not publishedAs labelled
Fugu Maxsakana/fugu-maxSakana AIWorld Cup opener draw, Dec 2025not publishedAs labelled
Fugu Ultrasakana/fugu-ultraSakana AI2025 Nobel Peace Prize, Oct 2025not publishedAs labelled
Fugu Ultra v2sakana/fugu-ultra-v2Sakana AIGLM-5 release, Feb 2026Aug 2026As labelled
Gemini 2.5 Flashgoogle/gemini-2.5-flashnot reportedJul 2024: Llama 3.1 405BJan 2025As labelled
Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-litenot reportedMay 2024: GPT-4oJan 2025As labelled
Gemini 2.5 Progoogle/gemini-2.5-pronot reportedNov 2024: Trump won US electionJan 2025As labelled
Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-previewGoogle, supplier's own keyNov 2024: Trump won US electionJan 2025As labelled
Gemini 3 Flash Previewgoogle/gemini-3-flash-previewnot reportedJan 2025: DeepSeek-R1Jan 2025As labelled
Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-litenot reportedFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-previewGoogle AI StudioFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-previewnot reportedFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtoolsGoogle AI StudioFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemini 3.5 Flashgoogle/gemini-3.5-flashnot reportedFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-litenot reportedFeb 2026: Norway top at Winter OlympicsMar 2026As labelled
Gemini 3.6 Flashgoogle/gemini-3.6-flashnot reportedFeb 2026: Norway top at Winter OlympicsMar 2026As labelled
Gemini 3.7 Flashgemini-3.7-flashnot reportedFeb 2025: Claude 3.7 SonnetMar 2026Cannot tell yetKnew events to Feb 2025; its maker says some topics stop at Jan 2025, so this may be normal for it.
Gemini 3.8 Flashgoogle/gemini-3.8-flashnot reportedFeb 2026: Norway top at Winter OlympicsMar 2026As labelled
Gemini Flash Latest~google/gemini-flash-latestGoogleClaude 3.7 Sonnet, Feb 2025Mar 2026Cannot tell yetServed by Google itself; it knew little after early 2025, which Google's card allows for some topics.
Gemini Pro Latest~google/gemini-pro-latestGoogle, supplier's own keyFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemma 2 27Bgoogle/gemma-2-27b-itNextBitNov 2023: OpenAI DevDay, GPT-4 Turbonot publishedAs labelled
Gemma 3 12Bgoogle/gemma-3-12b-itDeepInfraMay 2024: GPT-4oAug 2024As labelled
Gemma 3 27Bgoogle/gemma-3-27b-itParasailMay 2024: GPT-4oAug 2024As labelled
Gemma 3 4Bgoogle/gemma-3-4b-itDeepInfrano answer in this testAug 2024ConsistentNo answer: the model was rate-limited during this test.
Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-itDarkbloomno answer in this testJan 2025As labelled
Gemma 4 26B A4B Uncensoredvenice/e2ee-gemma-4-26b-a4b-uncensored-pnot reportedClaude 3.7 Sonnet, Feb 2025Jan 2025As labelled
Gemma 4 31Bgoogle/gemma-4-31b-itDeepInfraFeb 2025: Claude 3.7 SonnetJan 2025As labelled
Gemma 4 31B (Private via TEE)private/gemma4-31bnot reportedClaude 3.7 Sonnet, Feb 2025Jan 2025As labelled
Gemma 4 Uncensoredvenice/gemma-4-uncensorednot reportedDeepSeek-R1, Jan 2025Jan 2025As labelled
GLM 4.5z-ai/glm-4.5Z.AIDeepSeek-V3, Dec 2024not publishedConsistentIt ran out of room before answering this test; an earlier check fitted its expected age.
GLM 4.5 Airz-ai/glm-4.5-airNovitao1-preview, Sep 2024not publishedAs labelled
GLM 4.5Vz-ai/glm-4.5vNovitaDeepSeek-R1, Jan 2025not publishedAs labelled
GLM 4.6z-ai/glm-4.6Z.AIDeepSeek-R1, Jan 2025not publishedAs labelled
GLM 4.6Vz-ai/glm-4.6vZ.AIo1-preview, Sep 2024not publishedAs labelled
GLM 4.7z-ai/glm-4.7DeepInfraGemini 2.0 Flash, Dec 2024not publishedAs labelled
GLM 4.7 Flashz-ai/glm-4.7-flashCloudflareUS 'Liberation Day' tariffs, Apr 2025not publishedAs labelled
GLM 4.7 Flash Hereticvenice/olafangensan-glm-4.7-flash-hereticnot reportedPope Leo XIV elected, May 2025not publishedAs labelled
GLM 5z-ai/glm-5Amazon Bedrock, supplier's own keyClaude Opus 4, May 2025not publishedAs labelled
GLM 5 Turboz-ai/glm-5-turboZ.AIPope Leo XIV elected, May 2025not publishedAs labelled
GLM 5.1z-ai/glm-5.1PhalaClaude 4 models, May 2025not publishedAs labelled
GLM 5.2z-ai/glm-5.2not reportedGPT-5, Aug 2025not publishedAs labelled
GLM 5.2 (Fast)glm-5.2-fastFireworks, supplier's own keyGPT-5, Aug 2025not publishedAs labelled
GLM 5.3glm-5.3not reportedClaude Opus 4.5, Nov 2025not publishedAs labelled
GLM 5.3 Flashz-ai/glm-5.3-flashnot reportedGemini 3, Nov 2025not publishedAs labelled
GLM 5.3 FlashXz-ai/glm-5.3-flashxZ.AIGemini 3 and Opus 4.5, Nov 2025not publishedAs labelled
GLM 5.3 Primez-ai/glm-5.3-primeAlibabaClaude Opus 4.5, Nov 2025not publishedAs labelled
GLM 5V Turboz-ai/glm-5v-turboZ.AIClaude 3.7 Sonnet, Feb 2025not publishedCannot tell yetRecalls events only into early 2025, well before its late-2025 base; served by its maker, so likely genuine but thin on recent news.
GLM Flash Latest~z-ai/glm-flash-latestTogetherGemini 3, Nov 2025not publishedAs labelled
GLM Latest~z-ai/glm-latestSail ResearchUS shutdown ended, Nov 2025not publishedAs labelled
glm-5.2 by Engyengy/glm-5.2not reportedGPT-5, Aug 2025not publishedAs labelled
GLM-5.3 (Private via TEE)private/glm-5-3not reportedNYC mayoral result, Nov 2025not publishedAs labelled
glm-5.3 by Engyengy/glm-5.3not reportedGemini 3 and Grok 4.1, Nov 2025not publishedAs labelled
GLM-5.3 Flash (Private via TEE)private/glm-5-3-flashnot reportedNYC mayoral result, Nov 2025not publishedAs labelled
glm-5.3-flash by Engyengy/glm-5.3-flashnot reportedGPT-5.1 and Gemini 3 (Nov 2025)not publishedAs labelled
GPT Astra Latest~openai/gpt-astra-latestAmazon Bedrock, supplier's own keyGemini 3.1 Pro, Feb 2026Apr 2026As labelled
GPT Audioopenai/gpt-audionot reportedno answerOct 2023Not testableAudio-only model: it rejects text-only questions, so the text quiz cannot test it.
GPT Audio Miniopenai/gpt-audio-mininot reportedno answerOct 2023Not testableAudio-only model: it rejects text-only questions, so the text quiz cannot test it.
GPT Chat Latestopenai/gpt-chat-latestOpenAIClaude Opus 4.8, May 2026Aug 2025As labelled
GPT Luna Latest~openai/gpt-luna-latestAmazon Bedrock, supplier's own keyWinter Olympics medal table, Feb 2026May 2026As labelled
GPT Mini Latest~openai/gpt-mini-latestnot reportedno answerAug 2025Not answeringThe request was refused upstream on both tries, so no quiz ran.
GPT Sol Latest~openai/gpt-sol-latestAmazon Bedrock, supplier's own keyWinter Olympics medal table, Feb 2026Apr 2026As labelled
GPT Terra Latest~openai/gpt-terra-latestAmazon Bedrock, supplier's own keyWorld Cup draw, Dec 2025Feb 2026As labelled
GPT-3.5 Turboopenai/gpt-3.5-turboOpenAINothing tested (all after its cutoff)Sep 2021ConsistentIts stated cutoff (Sep 2021) predates every event we test, so this check cannot tell; nothing contradicts the label.
GPT-3.5 Turbo (older v0613)openai/gpt-3.5-turbo-0613AzureOpenAI DevDay, GPT-4 Turbo, Nov 2023Sep 2021Cannot tell yetAnswers like a newer model: it knows late-2023 events, two years past this model's Sep 2021 cutoff.
GPT-3.5 Turbo 16kopenai/gpt-3.5-turbo-16kOpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Sep 2021Cannot tell yetAnswers like a newer model: it knows late-2023 events, two years past this model's Sep 2021 cutoff.
GPT-3.5 Turbo Instructopenai/gpt-3.5-turbo-instructOpenAINothing tested (all after its cutoff)Sep 2021ConsistentIts stated cutoff (Sep 2021) predates every event we test, so this check cannot tell; nothing contradicts the label.
GPT-4openai/gpt-4AzureOpenAI DevDay, GPT-4 Turbo, Nov 2023Dec 2023As labelled
GPT-4 Turboopenai/gpt-4-turboOpenAIGPT-4 launch, Mar 2023Dec 2023As labelled
GPT-4.1openai/gpt-4.1OpenAIGPT-4o, May 2024Jun 2024As labelled
GPT-4.1 Miniopenai/gpt-4.1-miniOpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Jun 2024As labelled
GPT-4.1 Nanoopenai/gpt-4.1-nanoOpenAIOpenAI DevDay, Nov 2023 (date off)Jun 2024Cannot tell yetThis tiny model gave several confident wrong answers, so the test could not confirm its knowledge either way.
GPT-4oopenai/gpt-4oOpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Oct 2023As labelled
GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13OpenAIClaude 3 family, Mar 2024Oct 2023As labelled
GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06AzureClaude 3 family, Mar 2024Oct 2023As labelled
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20OpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Oct 2023As labelled
GPT-4o-miniopenai/gpt-4o-miniOpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Oct 2023As labelled
GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18OpenAIOpenAI DevDay, GPT-4 Turbo, Nov 2023Oct 2023As labelled
GPT-5openai/gpt-5OpenAIOpenAI o1-preview, Sep 2024Sep 2024As labelled
GPT-5 Miniopenai/gpt-5-miniOpenAIClaude 3 and GPT-4o, 2024 (dates off)May 2024As labelled
GPT-5 Nanoopenai/gpt-5-nanoOpenAIClaude 3 family, Mar 2024May 2024As labelled
GPT-5 Proopenai/gpt-5-proOpenAINo answerSep 2024Not answeringIt used its whole output budget on reasoning and returned no answer, so we could not test it.
GPT-5.1openai/gpt-5.1OpenAIOpenAI o1-preview, Sep 2024Sep 2024As labelled
GPT-5.1-Codexopenai/gpt-5.1-codexAzureLlama 3.1 405B, Jul 2024Sep 2024As labelled
GPT-5.1-Codex-Maxopenai/gpt-5.1-codex-maxAzureGPT-4o, May 2024Sep 2024As labelled
GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-miniAzureGPT-4o, 2024 (dates off)Sep 2024As labelled
GPT-5.2openai/gpt-5.2OpenAIDeepSeek-R1, Jan 2025Aug 2025As labelled
GPT-5.2 Chatopenai/gpt-5.2-chatnot reportedno answerAug 2025Not answeringThe model could not be reached: the maker has retired it, so the test could not run.
GPT-5.2 Proopenai/gpt-5.2-proOpenAIDeepSeek-R1, Jan 2025Aug 2025As labelled
GPT-5.2-Codexopenai/gpt-5.2-codexAzureGPT-4o, May 2024Aug 2025Cannot tell yetOnly showed knowledge up to May 2024, over a year before its stated cutoff; its sibling models are also very cautious, so unproven.
GPT-5.3-Codexgpt-5.3-codexOpenAIGPT-4.1, Apr 2025Aug 2025As labelled
GPT-5.4openai/gpt-5.4Amazon Bedrock, supplier's own keyGrok 4 release, Jul 2025Aug 2025As labelled
GPT-5.4 Minigpt-5.4-miniOpenAIGPT-4.1, Apr 2025Aug 2025As labelled
GPT-5.4 Nanogpt-5.4-nanoOpenAIDeepSeek-R1, Jan 2025Aug 2025As labelled
GPT-5.4 Proopenai/gpt-5.4-proOpenAIClaude Opus 4, May 2025Aug 2025As labelled
GPT-5.5openai/gpt-5.5Amazon Bedrock, supplier's own keyClaude Opus 4.1, Aug 2025Dec 2025As labelled
GPT-5.5 Proopenai/gpt-5.5-proOpenAIno answer returnedDec 2025Not answeringIt spent its whole answer budget thinking and returned no text, so it could not be tested.
GPT-5.6 Lunaopenai/gpt-5.6-lunaAmazon Bedrock, supplier's own keyGPT-5.1, Nov 2025Feb 2026As labelled
GPT-5.6 Luna Proopenai/gpt-5.6-luna-proOpenAIGPT-5.1, Nov 2025Feb 2026As labelled
GPT-5.6 Solopenai/gpt-5.6-solAmazon Bedrock, supplier's own keyWorld Cup draw, Dec 2025Feb 2026As labelled
GPT-5.6 Sol Proopenai/gpt-5.6-sol-proOpenAIMamdani wins NYC, Nov 2025Feb 2026As labelled
GPT-5.6 Terraopenai/gpt-5.6-terraAmazon Bedrock, supplier's own keyGemini 3, Nov 2025Feb 2026As labelled
GPT-5.6 Terra Proopenai/gpt-5.6-terra-proOpenAIWorld Cup draw, Dec 2025Feb 2026As labelled
GPT-6 Astraopenai/gpt-6-astraAmazon Bedrock, supplier's own keyGemini 3.1 Pro, Feb 2026Apr 2026As labelled
GPT-6 Astra Progpt-6-astra-proOpenAIGemini 3.1 Pro, Feb 2026Apr 2026As labelled
GPT-6 Lunaopenai/gpt-6-lunaAmazon Bedrock, supplier's own keyNorway tops Olympics, Feb 2026May 2026As labelled
GPT-6 Luna Proopenai/gpt-6-luna-proOpenAIClaude Opus 4.5, Nov 2025May 2026As labelled
GPT-6 Solgpt-6-solAmazon Bedrock, supplier's own keyGemini 3 and NYC mayor result, Nov 2025Apr 2026As labelled
GPT-6 Sol Proopenai/gpt-6-sol-proOpenAIGemini 3.1 Pro, Feb 2026Apr 2026As labelled
GPT-OSS 120B (Private via TEE)private/gpt-oss-120bnot reportedGPT-4o, May 2024Jun 2024As labelled
gpt-oss-120bopenai/gpt-oss-120bAmazon Bedrock, supplier's own keyGPT-4o, May 2024Jun 2024As labelled
gpt-oss-20bopenai/gpt-oss-20bAmazon Bedrock, supplier's own keyGPT-4o, 2024Jun 2024As labelled
gpt-oss-safeguard-20bopenai/gpt-oss-safeguard-20bGroqGPT-4o, May 2024not publishedNot testableA safety classifier, not a chat model, so a knowledge quiz does not apply; its answers match its base model.
Granite 4.0 Microibm-granite/granite-4.0-h-microCloudflarenone reliable (dates all wrong)not publishedCannot tell yetA very small model that gave wrong dates even for 2023 events, so its knowledge could not be pinned down.
Granite 4.2 8Bibm-granite/granite-4.2-8bDeepInfraMay 2024: GPT-4onot publishedCannot tell yetKnew events to about May 2024; its maker publishes no cutoff, and a sibling model looks the same.
Grok 4.20x-ai/grok-4.20xAIDeepSeek-R1, Jan 2025not publishedAs labelled
Grok 4.20 Multi-Agentx-ai/grok-4.20-multi-agentxAIPope Leo XIV elected, May 2025not publishedAs labelled
Grok 4.3x-ai/grok-4.3xAIDeepSeek-R1, Jan 2025not publishedCannot tell yetKnows events only to Jan 2025, older than a reported Dec 2025 cutoff; xAI has not published one, so we can't confirm.
Grok 4.5x-ai/grok-4.5xAIClaude Opus 4.6, Feb 2026not publishedAs labelled
Grok 4.6grok-4.6Amazon Bedrock, supplier's own keyGemini 3 launch, Nov 2025not publishedAs labelled
Grok 4.7x-ai/grok-4.7xAIGemini 3.1 Pro, Feb 2026May 2026As labelled
Grok Build 0.1x-ai/grok-build-0.1xAIPope Leo XIV elected, May 2025not publishedAs labelled
Grok Latest~x-ai/grok-latestxAIGemini 3.1 Pro, Feb 2026May 2026As labelled
Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405bDeepInfraGPT-4 Turbo, Nov 2023Dec 2023As labelled
Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70bDeepInfraLlama 2, Jul 2023Dec 2023As labelled
Hermes 4 405Bnousresearch/hermes-4-405bNebiusLlama 3.1, Jul 2024Dec 2023As labelled
Hunyuan A13B Instructtencent/hunyuan-a13b-instructSiliconFlowLlama 3.1 405B, Jul 2024not publishedAs labelled
Hy-MT2-1.8Btencent/hy-mt2-1.8bTencentnone shownnot publishedNot testableTranslation-only model, so a general-knowledge test does not apply.
Hy-MT2-30B-A3Btencent/hy-mt2-30b-a3bTencentGPT-5 date, Aug 2025 (bare date only)not publishedNot testableTranslation-only model, so a general-knowledge test does not apply.
Hy-MT2-7Btencent/hy-mt2-7bTencentLlama 2, Jul 2023not publishedNot testableTranslation-only model, so a general-knowledge test does not apply.
Hy3tencent/hy3PhalaDeepSeek-V3, Dec 2024not publishedCannot tell yetKnew events up to Dec 2024; the maker gives no cutoff, so we cannot tell whether that is normal for it.
Hy3 previewtencent/hy3-previewGMICloudLlama 4, Apr 2025not publishedAs labelled
Hy4 previewtencent/hy4-previewNovitaGemini 3, Nov 2025not publishedAs labelled
Inklingthinkingmachines/inklingDeepInfraNorway tops Winter Olympics, Feb 2026not publishedAs labelled
Inkling Smallthinkingmachines/inkling-smallDeepInfraNYC mayoral result, Nov 2025not publishedAs labelled
KAT-Coder-Pro V2.5kwaipilot/kat-coder-pro-v2.5not reportednot testednot publishedNot answeringThe model returned an error, so it could not be tested.
Kimi K2 0711moonshotai/kimi-k2NovitaClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Kimi K2 0905moonshotai/kimi-k2-0905NovitaPope Leo XIV elected, May 2025not publishedAs labelled
Kimi K2 Thinkingmoonshotai/kimi-k2-thinkingNovitaLlama 3.1, Jul 2024not publishedCannot tell yetShowed older knowledge than expected for this model; not conclusive, so we are rechecking it.
Kimi K2.5moonshotai/kimi-k2.5Amazon Bedrock, supplier's own keyPope Leo XIV elected, May 2025not publishedAs labelled
Kimi K2.6moonshotai/kimi-k2.6DecartPope Leo XIV elected, May 2025not publishedAs labelled
Kimi K2.7 Codemoonshotai/kimi-k2.7-codeMoonshot AIPope Leo XIV elected, May 2025not publishedAs labelled
Kimi K3moonshotai/kimi-k3not reportedKimi K2 Thinking, Nov 2025not publishedAs labelled
Kimi K3 (Fast)kimi-k3-fastFireworks, supplier's own keyWinter Olympics medal table, Feb 2026not publishedAs labelled
Kimi K3 (Private via TEE)private/kimi-k3not reportedNYC mayoral result, Nov 2025not publishedAs labelled
Kimi Latest~moonshotai/kimi-latestFireworks, supplier's own keyKimi K2 Thinking, Nov 2025not publishedAs labelled
kimi-k3 by Engyengy/kimi-k3not reportedKimi K2 Thinking (Nov 2025)not publishedAs labelled
Laguna S 2.1poolside/laguna-s-2.1PoolsidePope Leo XIV (Prevost), May 2025not publishedAs labelled
Laguna XS 2.1poolside/laguna-xs-2.1PoolsideLlama 4 and Pope Leo XIV, 2025not publishedAs labelled
LFM2.5-2.6Bliquid/lfm-2.5-2.6bnot reportedno answernot publishedNot answeringWe could not run this model: no server was available to answer it.
Ling 3.0 Flashinclusionai/ling-3.0-flashDeepInfraDeepSeek-R1, Jan 2025not publishedAs labelled
Ling 3.0 Flash Fininclusionai/ling-3.0-flash-finDeepInfraUS tariffs, Apr 2025not publishedAs labelled
Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-santenot reportedno answernot publishedNot answeringWe could not run this model: no server was available to answer it.
Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vlNovitaPope Leo XIV elected, May 2025not publishedCannot tell yetKnowledge ends around mid-2025, a year before the maker's stated data date, probably because its text side is older.
Llama 3 8B Lunarissao10k/l3-lunaris-8bParasailnothing on the listMar 2023ConsistentBuilt on a model whose knowledge ends in early 2023, so recent events are out of its range.
Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instructAmazon Bedrock, supplier's own keyGPT-4 and Llama 2, 2023Dec 2023As labelled
Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instructDeepInfraGPT-4, Mar 2023Dec 2023As labelled
Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70bDeepInfraDevDay and GPT-4 Turbo, Nov 2023Dec 2023As labelled
Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instructCloudflarenothing reliableDec 2023Cannot tell yetA very small model; it guesses dates, so its answers about events cannot be trusted.
Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instructParasailOpenAI DevDay, Nov 2023Dec 2023As labelled
Llama 3.3 70B (Private via TEE)private/llama3-3-70bnot reportedLlama 2, Jul 2023Dec 2023As labelled
Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instructCloudflareGPT-4 and Llama 2, 2023Dec 2023As labelled
Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70bNextBitDevDay and GPT-4 Turbo, Nov 2023Dec 2023As labelled
Llama 4 Maverickmeta-llama/llama-4-maverickDeepInfraOpenAI o1-preview, Sep 2024Aug 2024As labelled
Llama 4 Scoutmeta-llama/llama-4-scoutDeepInfraLlama 3.1 405B, Jul 2024Aug 2024As labelled
Llama Guard 4 12Bmeta-llama/llama-guard-4-12bDeepInfradid not saynot publishedNot testableA safety classifier that only replies safe or unsafe, so its knowledge cannot be tested.
LongCat 2.0meituan/longcat-2.0AtlasCloudClaude Opus 4.5, Nov 2025not publishedAs labelled
Magnum v4 72Banthracite-org/magnum-v4-72bMancer 2GPT-4 Turbo, Nov 2023not publishedCannot tell yetIt gave the same made-up date for most questions, so we could not check its knowledge.
Mercury 2inception/mercury-2InceptionGPT-4o, May 2024not publishedCannot tell yetIts knowledge stops around mid-2024, early for a 2026 model; the maker publishes no cutoff, so we cannot confirm.
Mercury 2.5inception/mercury-2.5InceptionClaude 3.7 Sonnet, Feb 2025not publishedCannot tell yetIts knowledge stops around early 2025, early for a late-2026 model; the maker publishes no cutoff.
MiMo-V2.5xiaomi/mimo-v2.5XiaomiLouvre jewel theft, Oct 2025not publishedAs labelled
MiMo-V2.5-Proxiaomi/mimo-v2.5-proGMICloudPope Leo XIV elected, May 2025not publishedAs labelled
MiMo-V2.6-Flashxiaomi/mimo-v2.6-flashXiaomiWinter Olympics result, Feb 2026not publishedAs labelled
MiMo-V2.6-Proxiaomi/mimo-v2.6-proDeepInfraGemini 3, Nov 2025not publishedAs labelled
MiMo-V2.6-Pro-UltraSpeedxiaomi/mimo-v2.6-pro-ultraspeedXiaomiGPT-5, Aug 2025not publishedAs labelled
MiniMax M1minimax/minimax-m1MinimaxClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
MiniMax M2minimax/minimax-m2MinimaxClaude 3 family, Mar 2024not publishedCannot tell yetOn this test it said it knew nothing after early 2024, but it was very cautious; a follow-up is needed before drawing conclusions.
MiniMax M2-herminimax/minimax-m2-herMinimaxPope Leo XIV elected, May 2025not publishedAs labelled
MiniMax M2.1minimax/minimax-m2.1MinimaxLlama 3.1 405B, Jul 2024not publishedCannot tell yetShowed nothing after mid-2024, even on MiniMax's own servers; MiniMax publishes no cutoff to check it against.
MiniMax M2.5minimax/minimax-m2.5FriendliClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
MiniMax M2.7minimax/minimax-m2.7DeepInfraTrump election win, Nov 2024not publishedCannot tell yetShowed nothing after late 2024 and denied Pope Leo XIV exists; MiniMax publishes no cutoff to check it against.
MiniMax M3minimax/minimax-m3StreamLakeClaude Sonnet 4.5, Sep 2025not publishedAs labelled
MiniMax-01minimax/minimax-01MinimaxGPT-4o, May 2024not publishedAs labelled
Ministral 3 14B 2512mistralai/ministral-14b-2512MistralGPT-4o, May 2024not publishedAs labelled
Ministral 3 3B 2512mistralai/ministral-3b-2512MistralGPT-4o, May 2024not publishedAs labelled
Ministral 3 8B 2512mistralai/ministral-8b-2512MistralGPT-4o, May 2024not publishedAs labelled
Mistral Largemistralai/mistral-largeMistralGemini 1.5 Pro, Feb 2024 (earlier test)not publishedConsistentThe date test was rate-limited and got no answer; an earlier quiz showed early-2024 knowledge.
Mistral Large 2407mistralai/mistral-large-2407MistralDeepSeek-V3, Dec 2024Oct 2023As labelled
Mistral Large 3 2512mistralai/mistral-large-2512MistralGemini 1.5 Pro, Feb 2024 (earlier test)not publishedConsistentThe date test was rate-limited and got no answer; an earlier quiz showed only early-2024 knowledge, so it needs a retest.
Mistral Medium 3mistralai/mistral-medium-3MistralOpenAI o1-preview, Sep 2024not publishedAs labelled
Mistral Medium 3.1mistralai/mistral-medium-3.1MistralOpenAI o1-preview, Sep 2024not publishedAs labelled
Mistral Medium 3.5mistralai/mistral-medium-3-5MistralMistral Large 2, Jul 2024not publishedCannot tell yetServed by Mistral itself but showed nothing after mid-2024 for an April 2026 model; Mistral publishes no cutoff.
Mistral Nemomistralai/mistral-nemoDeepInfraLlama 2, Jul 2023not publishedAs labelled
Mistral Small 3mistralai/mistral-small-24b-instruct-2501DeepInfraClaude 3 models, Mar 2024Oct 2023As labelled
Mistral Small 3.1 24Bmistralai/mistral-small-3.1-24b-instructCloudflareClaude 3 models, Mar 2024Oct 2023As labelled
Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instructDeepInfraGPT-4o, May 2024Oct 2023As labelled
Mistral Small 4mistralai/mistral-small-2603MistralClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Mixtral 8x22B Instructmistralai/mixtral-8x22b-instructMistralnothing usable yetnot publishedConsistentThe date test was rate-limited and got no answer, so this model has not been checked yet.
Morph V3 Fastmorph/morph-v3-fastMorphdid not saynot publishedNot testableA code-editing model, not a chat model: it echoed the questions back, so a knowledge quiz does not apply.
Morph V3 Largemorph/morph-v3-largeMorphdid not saynot publishedNot testableA code-editing model, not a chat model: it did not engage with the questions, so a knowledge quiz does not apply.
Muse Glimmer 30Bmeta/muse-glimmer-30bTogetherGPT-5, Aug 2025Jan 2026As labelled
Muse Spark 1.1meta/muse-spark-1.1MetaMamdani elected NYC mayor, Nov 2025not publishedAs labelled
Muse Spark 1.2meta/muse-spark-1.2MetaMamdani elected NYC mayor, Nov 2025not publishedAs labelled
Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributornot reportedno answernot publishedNot answeringBlocked: this tier trains on your data, and our data policy stops requests from reaching it.
Muse Spark 1.3meta/muse-spark-1.3MetaMamdani elected NYC mayor, Nov 2025not publishedAs labelled
Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributornot reportedno answernot publishedNot answeringBlocked: this tier trains on your data, and our data policy stops requests from reaching it.
MythoMax 13Bgryphe/mythomax-l2-13bParasailnothing on this list (all 2023+)Sep 2022ConsistentBuilt on a 2022-era base model; too old for this test to check.
Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3bNebiuso1-preview, Sep 2024Jun 2025As labelled
Nemotron 3 Nano Omninvidia/nemotron-3-nano-omni-30b-a3b-reasoningnot reportedno answernot publishedNot answeringNo paid route for this model id was available when we tested, so the quiz could not run.
Nemotron 3 Supernvidia/nemotron-3-super-120b-a12bDekaLLMGPT-4o, May 2024Jun 2025Cannot tell yetShowed about a year less knowledge than its maker states; not conclusive, so we are rechecking it.
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55bBaseTenDeepSeek-R1, Jan 2025Sep 2025As labelled
Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safetyDeepInfradid not saynot publishedNot testableA safety classifier, not a chat model: it only labels messages safe or unsafe, so a knowledge quiz does not apply.
Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightningIo Neto1-preview, Sep 2024Sep 2025Cannot tell yetShowed about a year less knowledge than its maker states; not conclusive, so we are rechecking it.
North Mini Codecohere/north-mini-codenot reportedno answernot publishedNot answeringNo paid version of this model is being served, so the request failed; it should be removed from the list.
Nova 2 Liteamazon/nova-2-lite-v1Amazon Bedrock, supplier's own keyDeepSeek V2, May 2024Oct 2025Cannot tell yetKnew nothing after mid-2024 and invented later events, though Amazon states an Oct 2025 cutoff; may be an older Nova.
Nova Lite 1.0amazon/nova-lite-v1Amazon Bedrock, supplier's own keyGPT-4o, May 2024Oct 2024As labelled
Nova Micro 1.0amazon/nova-micro-v1Amazon Bedrock, supplier's own keyLlama 2, Jul 2023Oct 2024Cannot tell yetIts answers were thin and partly wrong, so we could not confirm how recent its knowledge is.
Nova Premier 1.0amazon/nova-premier-v1not reportedno answerOct 2024Not answeringThe model has been retired by its maker and returned an error; it should be removed from the list.
Nova Pro 1.0amazon/nova-pro-v1Amazon Bedrock, supplier's own keyGPT-4 Turbo, Nov 2023Oct 2024Cannot tell yetIt recalled events only up to late 2023, about 11 months before Amazon's stated cutoff.
o1openai/o1OpenAIGPT-4 Turbo, Nov 2023Oct 2023As labelled
o1-proopenai/o1-proOpenAIGPT-4 Turbo, Nov 2023Oct 2023As labelled
o3openai/o3OpenAIGPT-4o, May 2024Jun 2024As labelled
o3 Miniopenai/o3-miniOpenAIGPT-4 Turbo, Nov 2023Oct 2023As labelled
o3 Mini Highopenai/o3-mini-highOpenAIGPT-4 Turbo, Nov 2023Oct 2023As labelled
o3 Proopenai/o3-proOpenAIGPT-4o, May 2024Jun 2024As labelled
o4 Miniopenai/o4-miniOpenAIGPT-4o, 2024Jun 2024As labelled
o4 Mini Highopenai/o4-mini-highOpenAIGPT-4o, 2024Jun 2024As labelled
Palmyra X5writer/palmyra-x5Amazon Bedrock, supplier's own keyLlama 3.1 405B, Jul 2024not publishedAs labelled
Paretounbiased/paretoUnbiasedWorld Cup opener pairing, Dec 2025not publishedAs labelled
Perceptron Mk1perceptron/perceptron-mk1PerceptronLlama 4 and Pope Leo XIV, 2025not publishedAs labelled
Perceptron Mk1.5perceptron/perceptron-mk1.5PerceptronMamdani NYC win, Nov 2025not publishedAs labelled
Phi 4microsoft/phi-4DeepInfraClaude 3 family, 2024Jun 2024As labelled
Qwen Plus 0728qwen/qwen-plus-2025-07-28Alibabao1-preview, Sep 2024not publishedAs labelled
Qwen-Plusqwen/qwen-plusAlibabao1-preview, Sep 2024not publishedAs labelled
Qwen2.5 72B Instructqwen/qwen-2.5-72b-instructDeepInfraOpenAI DevDay, late 2023not publishedAs labelled
Qwen2.5 7B Instructqwen/qwen-2.5-7b-instructPhalanothing shownnot publishedCannot tell yetRefused every question, even GPT-4, so this test says nothing about which model it is.
Qwen2.5 Coder 32B Instructqwen/qwen-2.5-coder-32b-instructCloudflareOpenAI DevDay, late 2023not publishedAs labelled
Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instructParasailGPT-4o, May 2024not publishedAs labelled
Qwen3 14Bqwen/qwen3-14bNextBitDeepSeek-V3, late 2024not publishedAs labelled
Qwen3 235B A22Bqwen/qwen3-235b-a22bAlibabaDeepSeek-V3, late 2024not publishedAs labelled
Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507DeepInfrao1-preview, 2024not publishedAs labelled
Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507NovitaLlama 4, Apr 2025not publishedAs labelled
Qwen3 30B A3Bqwen/qwen3-30b-a3bDeepInfraLlama 2, Jul 2023 (later answers vague)not publishedCannot tell yetAnswered with years only, often wrong, so this test cannot date it.
Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507StreamLakeUS election result, Nov 2024not publishedAs labelled
Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507AlibabaTrump wins US election, Nov 2024not publishedAs labelled
Qwen3 32Bqwen/qwen3-32bDeepInfrao1-preview, Sep 2024not publishedAs labelled
Qwen3 8Bqwen/qwen3-8bAlibabao1-preview, Sep 2024not publishedAs labelled
Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instructSiliconFlowno usable answernot publishedConsistentThis model gave no usable answer to our knowledge test, so we could not check it.
Qwen3 Coder 480B A35Bqwen/qwen3-coderDeepInfraDeepSeek-R1, Jan 2025not publishedAs labelled
Qwen3 Coder Flashqwen/qwen3-coder-flashAlibabaTrump beats Harris, Nov 2024not publishedAs labelled
Qwen3 Coder Nextqwen/qwen3-coder-nextAlibabaClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Qwen3 Coder Plusqwen/qwen3-coder-plusAlibabaTrump wins US election, Nov 2024not publishedAs labelled
Qwen3 Maxqwen/qwen3-maxAlibabaQwen3 release, Apr 2025not publishedAs labelled
Qwen3 Max Thinkingqwen/qwen3-max-thinkingAlibabaQwen3 release, Apr 2025not publishedAs labelled
Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instructAlibabao1-preview, Sep 2024not publishedAs labelled
Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinkingGoogleTrump wins US election, Nov 2024not publishedAs labelled
Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instructParasailo1-preview, Sep 2024not publishedAs labelled
Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinkingAlibabaClaude 3.5 Sonnet, Jun 2024not publishedAs labelled
Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instructDeepInfraTrump wins US election, Nov 2024not publishedAs labelled
Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinkingAlibabaRebels take Aleppo, Nov 2024not publishedAs labelled
Qwen3 VL 32B Instructqwen/qwen3-vl-32b-instructAlibabaTrump wins US election, Nov 2024not publishedAs labelled
Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instructParasailClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Qwen3 VL 8B Thinkingqwen/qwen3-vl-8b-thinkingAlibabaDeepSeek-R1, Jan 2025not publishedAs labelled
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17bPhalaLlama 3, Apr 2024not publishedCannot tell yetIts answers stop in early 2024, well before its likely cutoff. The maker's own version answers the same way.
Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15Alibabao1-preview, Sep 2024not publishedCannot tell yetIt showed knowledge only to late 2024; its maker publishes no cutoff, so we cannot say whether that is expected.
Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420AlibabaClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10bNovitaQwen2.5 release, Sep 2024not publishedCannot tell yetIts answers stop around 2024, well before its likely cutoff. The maker's own version answers the same way.
Qwen3.5-27Bqwen/qwen3.5-27bNovitaGPT-4o, May 2024not publishedCannot tell yetIts answers stop in mid-2024, well before its likely cutoff. The maker's own version answers the same way.
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3bDarkbloomGPT-4o, May 2024not publishedCannot tell yetIts answers stop in mid-2024, well before its likely cutoff. This comes straight from the maker, so it is how the model answers.
Qwen3.5-9Bqwen/qwen3.5-9bDarkbloomno answer givennot publishedNot answeringIt spent its whole answer budget thinking and gave no reply, so it could not be tested.
Qwen3.5-Flashqwen/qwen3.5-flash-02-23Alibabao1-preview, Sep 2024not publishedCannot tell yetIts answer was cut off mid-thought; it seemed to know events to late 2024, and its maker publishes no cutoff to compare.
Qwen3.6 27Bqwen/qwen3.6-27bChutesClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3bDarkbloomDeepSeek-R1, Jan 2025not publishedAs labelled
Qwen3.6 Flashqwen/qwen3.6-flashAlibabaDeepSeek-R1, Jan 2025not publishedAs labelled
Qwen3.6 Max Previewqwen/qwen3.6-max-previewAlibabaNYC mayoral result, Nov 2025not publishedAs labelled
Qwen3.6 Plusqwen/qwen3.6-plusAlibabaClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
qwen3.6-35b-a3b by Engyengy/qwen3.6-35b-a3bnot reportedClaude 3.7 Sonnet (Feb 2025)not publishedAs labelled
Qwen3.7 Flashqwen/qwen3.7-flashAlibabaDeepSeek-R1, Jan 2025not publishedCannot tell yetReleased Jul 2026 but showed nothing past Jan 2025; Qwen models under-report, so this is not proof of a swap.
Qwen3.7 Maxqwen/qwen3.7-maxAlibabaLlama 4, Apr 2025not publishedAs labelled
Qwen3.7 Plusqwen/qwen3.7-plusAlibabaGPT-4.5 and Grok 3, Feb 2025not publishedCannot tell yetReleased Jun 2026 but showed nothing past Feb 2025; Qwen models under-report, so this is not proof of a swap.
Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95bTogetherNYC mayoral result, Nov 2025not publishedAs labelled
Qwen3.8 27Bqwen/qwen3.8-27bNovitaGemini 3 launch, Nov 2025not publishedAs labelled
Qwen3.8 Flashqwen/qwen3.8-flashAlibabano answer givennot publishedNot answeringIt spent its whole answer budget thinking and gave no reply, so it could not be tested.
Qwen3.8 Max (0902)qwen/qwen3.8-max-0902not reportedPope Leo XIV, May 2025not publishedAs labelled
Qwen3.8 Max Primeqwen/qwen3.8-max-primeAlibabaPope Leo XIV, May 2025not publishedAs labelled
Qwen3.8 Omni Flashqwen/qwen3.8-omni-flashAlibabaGrok 4.1, Nov 2025not publishedAs labelled
qwen3.8-27b by Engyengy/qwen3.8-27bnot reportedGPT-5 (Aug 2025)not publishedAs labelled
R1deepseek/deepseek-r1Novitanot measurednot publishedConsistentThis model gave no usable answer to our knowledge test, so we could not check it.
R1 0528deepseek/deepseek-r1-0528Novitanot measurednot publishedConsistentThis model gave no usable answer to our knowledge test, so we could not check it.
R1 Distill Llama 70Bdeepseek/deepseek-r1-distill-llama-70bNovitaGPT-4o, May 2024Dec 2023As labelled
Reka Edgerekaai/reka-edgeRekanothing it could date correctlynot publishedCannot tell yetThis small model gave muddled answers, so its knowledge could not be dated.
Reka Flash 3rekaai/reka-flash-3RekaGPT-4, Mar 2023not publishedCannot tell yetIts answer was cut off after one garbled item, so its knowledge could not be dated.
Relace Apply 3relace/relace-apply-3not reportedno answernot publishedNot testableA code-merge tool that only accepts code edits, so it cannot take a knowledge quiz; it needs a merge test.
Relace Searchrelace/relace-searchRelace2024 US election, Nov 2024not publishedNot testableA code-search tool rather than a chat model, so a general-knowledge test does not apply.
ReMM SLERP 13Bundi95/remm-slerp-l2-13bMancer 2nothing on the ladderSep 2022ConsistentBuilt on a 2022-era base, so it cannot know anything this test asks about.
Sabamistralai/mistral-sabanot reportedClaude 3.7 Sonnet, Feb 2025not publishedAs labelled
Sakana Namazusakana/sakana-namazunot reportedno answernot publishedNot answeringNo server could run it under our privacy settings, so we got no answer to test.
Schematron V2 Smallinference-net/schematron-v2-smallInferenceNetdid not saynot publishedNot testableAn extraction model that turns text into JSON; it does not answer questions, so its knowledge cannot be tested.
Schematron V2 Turboinference-net/schematron-v2-turboInferenceNetdid not saynot publishedNot testableAn extraction model that turns text into JSON; it does not answer questions, so its knowledge cannot be tested.
Seed 1.6bytedance-seed/seed-1.6SeedLlama 3.1 405B (Jul 2024)not publishedAs labelled
Seed 1.6 Flashbytedance-seed/seed-1.6-flashSeedLlama 3.1 (2024, month unsure)not publishedCannot tell yetIts dates for well-known 2023 to 2024 events were mostly wrong, so its knowledge horizon could not be measured.
Seed 2.1 Turbobytedance-seed/seed-2-1-turboSeedClaude 3.5 Sonnet update (Oct 2024)not publishedCannot tell yetIts answers suggest knowledge ending in 2024, well before this 2026 model's release; not confirmed.
Seed-2.0-Codebytedance-seed/seed-2.0-codeSeedOpenAI o1-preview (Sep 2024)not publishedCannot tell yetIts knowledge appears to end in 2024, over a year before this model's 2026 release; the cause is unconfirmed.
Seed-2.0-Litebytedance-seed/seed-2.0-liteSeedLlama 3.1 405B (Jul 2024)not publishedCannot tell yetIts knowledge appears to end mid-2024, well before this model's 2026 release; the cause is unconfirmed.
Seed-2.0-Minibytedance-seed/seed-2.0-miniSeedOpenAI o1-preview (Sep 2024)not publishedCannot tell yetIts answers were inconsistent and point to knowledge ending in 2024; the cause is unconfirmed.
Skyfall 36B V2thedrummer/skyfall-36b-v2ParasailClaude 3 family, Mar 2024not publishedAs labelled
Solar Mini 4upstage/solar-mini4UpstagePope Leo XIV elected, May 2025not publishedCannot tell yetKnew events to about May 2025, about 10 months short of what its release date suggests; not enough to call it a swap.
Solar Pro 3upstage/solar-pro-3Upstageo1-preview, Sep 2024not publishedCannot tell yetRecalled events only to about late 2024 and got some wrong; served by the maker, so likely its own limit.
Solar Pro 4upstage/solar-pro4UpstageClaude Sonnet 4, May 2025not publishedAs labelled
Sonarperplexity/sonarPerplexityClaude 3.7 Sonnet, Feb 2025not publishedNot testableAnswers by searching the web, so a knowledge-date test does not apply.
Sonar Deep Researchperplexity/sonar-deep-researchPerplexityDeepSeek-R1, Jan 2025 (via search)not publishedNot testableAnswers by searching the web, so a knowledge-date test does not apply.
Sonar Properplexity/sonar-proPerplexityGPT-5, Aug 2025not publishedNot testableAnswers by searching the web, so a knowledge-date test does not apply.
Sonar Pro Searchperplexity/sonar-pro-searchPerplexityWorld Cup opener, Jun 2026not publishedNot testableAnswers by searching the web, so a knowledge-date test does not apply.
Sonar Reasoning Properplexity/sonar-reasoning-proPerplexityo1-preview, Sep 2024not publishedNot testableAnswers by searching the web, so a knowledge-date test does not apply.
Step 3.5 Flashstepfun/step-3.5-flashSiliconFlowo1-preview, Sep 2024not publishedCannot tell yetIt showed knowledge only to late 2024; its maker publishes no cutoff, and its sibling on the maker's own service answers the same.
Step 3.7 Flashstepfun/step-3.7-flashStepFunClaude 3.5 Sonnet, Jun 2024not publishedCannot tell yetIt showed knowledge only to mid 2024, well short of its 2026 release; its maker's own service answers the same way.
Ternary Bonsai 2 27Bprism-ml/ternary-bonsai-2-27bDarkbloomGemini 3, late 2025not publishedAs labelled
Trinity Large Thinkingarcee-ai/trinity-large-thinkingArcee AIDeepSeek-R1 (Jan 2025)not publishedCannot tell yetIts knowledge seems to end around early 2025; the maker publishes no cutoff to compare against.
UI-TARS 7Bbytedance/ui-tars-1.5-7bParasailLlama 3.1 405B (Jul 2024)not publishedAs labelled
UnslopNemo 12Bthedrummer/unslopnemo-12bParasailLlama 2, Jul 2023not publishedAs labelled
Venice Role Play Uncensoredvenice/venice-uncensored-role-playnot reportedLlama 3.1 405B, mid 2024not publishedAs labelled
Venice Uncensored 1.2venice/venice-uncensored-1-2not reportedLlama 3.1 405B, Jul 2024not publishedAs labelled
Venice: Uncensoredcognitivecomputations/dolphin-mistral-24b-venice-editionVenice, supplier's own keyLlama 3.1 405B, Jul 2024Oct 2023As labelled
Voxtral Small 24B 2507mistralai/voxtral-small-24b-2507MistralOpenAI o1-preview, Sep 2024not publishedAs labelled
Weaver (alpha)mancer/weaverMancer 2nothing reliablenot publishedConsistentA 2023 role-play model; it invents answers about recent events, so do not rely on it for facts.
WizardLM-2 8x22Bmicrosoft/wizardlm-2-8x22bNovitaOpenAI o1-preview, Sep 2024not publishedCannot tell yetIt knows events from after this model came out in April 2024, so a newer, different model may be answering.

A note on the method's limits. Genuine models are cautious about the last months before their cutoff and often under-report what they know, so a short gap proves nothing and a long one is only a reason to look harder. Where a maker publishes no cutoff we use its release date minus six months. A hidden prompt can make a genuine model answer like an older one, which is why we test with the date and look at the route as well as the answers. We would rather say "cannot tell yet" than guess in either direction.

5. This is why we are already building our own router

Until this week, almost everything on askr ran through one supplier. That let us launch with hundreds of models in a few weeks. It also meant that what reached the model depended on choices made two layers below us, which we could not see and did not check. The instructions in front of Grok 4.7 are one of those choices. That is the bug behind the bug.

The fix is a router of our own: askr connects directly to the companies that make the models, one by one, and the supplier keeps only the long tail, behind the nightly check. We did not start this because of the audit. We already have one direct connection live: Claude (direct), a straight line to Anthropic's API, has been running since 25 September, and every Claude answer on it comes from Anthropic and says so. Grok comes next, directly from xAI, then GLM from Z.ai, then OpenAI, Google and DeepSeek. Each direct connection shows "(direct)" after the model's name and charges the maker's list price.

The rule from now on is short. A model on askr is the model on the label, served by its maker or by a route we name on the receipt, and checked every night.

Live

Claude (direct) to Anthropic, since 25 September. GPT (direct) to OpenAI, since 27 September. Grok (direct) to xAI, since 28 September.

Next

GLM (direct) to Z.ai.

Then

Google and DeepSeek direct. The supplier for the tail only, behind the nightly check.

6. What askr is

askr is not an inference market. We are building a consumer app, and it is already live: one place to talk to every major model, with Skills that turn a chat into a job done, Collab rooms where a group works with the models together on a shared balance, and Mind, a memory that follows you across all of it, coming soon.

The models are the ingredients. We do not make them and we do not pretend to. What we owe you is that the ingredients are what the label says, that you can see where they came from, and that when something looks wrong, you hear the whole story from us, whichever way it comes out. This page is that story. The nightly check keeps it true from now on.

7. Check it yourself

Here is how to test whether any model is the model on its label, the way we test ours now. Then the accusations made about askr, the proper test for each and what it showed, and why the original tests fell short.

How to test any model properly

  1. Tell it the date. Start with "Today's date is ...". Without it, a model assumes it is still early in its training and treats anything later as not yet happened. In our tests this one line turned Grok 4.7's "there was no papal election in May 2025" into "Leo XIV, elected 8 May 2025".
  2. Ask about the world, not about the model. Asking a model who it is proves nothing. Claude Haiku 4.5, taken straight from Anthropic, calls itself Claude 3.5 Sonnet.
  3. Ask for things that cannot be guessed. A winner's name, an exact release date. Skip anything announced in advance, such as where the Olympics are held or when the World Cup starts.
  4. Find the newest thing it gets right. Ask across two or three years of events. The newest one it gets right, with the right name and month, is where its knowledge ends. Ask the last few one at a time: in a long list, careful models say "never heard of it" too early.
  5. Compare with the maker's stated cutoff, and expect a gap. Genuine Claude, straight from Anthropic, knows events up to 3 to 8 months before its stated cutoff and gets vaguer after that. Only a gap of a year or more is a reason to look harder.
  6. Ask the maker's own copy the same way. Same date line, same questions, one at a time, with no other app in between. Official apps and coding assistants add instructions of their own, and some search the web. A difference between two like-for-like runs is the real signal.
  7. Make sure it is not searching the web. A model that searches knows yesterday's news and cites sources, so it tells you nothing about its training. Our table marks those "not testable".
  8. Count what you pay for. A bare "hi" should cost a handful of prompt tokens. Hundreds or thousands mean something sits in front of your message.

The ladder we used, short enough to paste into any chat. Add today's date in the first line.

Today's date is [today's date]. For each item, say in one line what it is and when it happened, month and year, from your own training knowledge. Try every item. If you have genuinely never heard of one, say so.
1) OpenAI's GPT-4o
2) The winner of the 2024 US presidential election
3) DeepSeek-R1
4) The pope elected in May 2025
5) OpenAI's GPT-5
6) The winner of the 2025 Nobel Peace Prize
7) Google's Gemini 3
8) The winner of the November 2025 New York City mayoral election
9) Anthropic's Claude Opus 4.6
10) The country that topped the 2026 Winter Olympics medal table
11) Google's Gemini 3.1 Pro
12) Anthropic's Claude Opus 4.8
ItemA correct answer names
1) OpenAI's GPT-4oMay 2024
2) The winner of the 2024 US presidential electionDonald Trump, November 2024
3) DeepSeek-R1January 2025
4) The pope elected in May 2025Leo XIV (Robert Prevost), 8 May 2025
5) OpenAI's GPT-57 August 2025
6) The winner of the 2025 Nobel Peace PrizeMaría Corina Machado, 10 October 2025
7) Google's Gemini 318 November 2025
8) The winner of the November 2025 New York City mayoral electionZohran Mamdani, 4 November 2025
9) Anthropic's Claude Opus 4.65 February 2026
10) The country that topped the 2026 Winter Olympics medal tableNorway, February 2026
11) Google's Gemini 3.1 Pro19 February 2026
12) Anthropic's Claude Opus 4.828 May 2026

Norway also topped the 2018 and 2022 medal tables, so count item 10 only alongside other 2026 answers. The newest item a model gets right is where its knowledge ends; set that against its maker's stated cutoff.

The accusations, and the proper test for each

"Grok 4.7 is really an older Grok, with knowledge to mid 2025"

Did not hold up
The proper test
Give it today's date, ask about early-2026 events one question at a time, and ask an older genuine model the same as a control.
What we found
It named GLM-5 (11 February 2026), Gemini 3.1 Pro (19 February 2026), Claude Opus 4.6 (5 February 2026) and Norway topping the 2026 medal table. Claude Haiku 4.5 from Anthropic, whose knowledge ends in early 2025, got the same line and could name none of them.
Why the original fell short
No date, so late-2025 events looked unconfirmed to the model. Most questions were about rivals' launches, which Grok models often deny: Grok 4.5 said Gemini 3.1 Pro does not exist while naming GLM-5. The official copy answered inside Cursor, which adds instructions of its own.

"Grok 4.6 is Grok 3, and its hidden script admits a 2024-10 cutoff"

Did not hold up
The proper test
Count the prompt tokens on a bare "hi" to see whether anything sits in front of it, then ask dated questions.
What we found
A bare "hi" counts 19 prompt tokens: nothing of that size is in front of it. It described the 19 October 2025 Louvre theft in detail and, with the date, named María Corina Machado's Nobel Peace Prize (10 October 2025). Grok 3 came out in February 2025 and cannot know either.
Why the original fell short
A model asked to repeat instructions it does not have will write out the kind it was trained with. "January 2026 has not yet occurred" is exactly what any model says when nobody tells it the date.

"GLM 5.3 is really GLM-4.6"

Did not hold up
The proper test
Ask, with the date, about events after GLM-4.6 came out on 30 September 2025, and put the same family questions to Z.ai's own model.
What we found
It named Gemini 3 (18 November 2025), Grok 4.1 (17 November 2025) and the 2025 Nobel Peace Prize winner. It does not know GLM-5 or GLM-4.7, and neither does GLM 5.3 FlashX served by Z.ai itself, asked the same way.
Why the original fell short
The family-tree questions hit a blind spot the maker's own model shares. The "real" GLM 5.3 answered inside a coding assistant on Z.ai's coding plan, which brings instructions of its own.

"It names itself three different ways"

Did not hold up
The proper test
Do not use self-identification.
What we found
Genuine models do the same: Claude Haiku 4.5 from Anthropic says it is Claude 3.5 Sonnet, and Claude Opus 5 says it is Opus 4.5.
Why the original fell short
A model's name for itself is not evidence of anything.

"About 1,250 hidden words are billed on every Grok 4.7 message"

Right, and refunded
The proper test
Send "hi" straight to our supplier, askr's code out of the path, and read the prompt token count.
What we found
1,243 prompt tokens, 1,152 of them cached, on the Grok 4.7 route, against 19 on Grok 4.6. askr does not add them. We have refunded them and asked our supplier where they come from (section 3).
Why the original fell short
It held. Credit to the user.

"'The conversation is too long' comes from Grok's consumer website"

Our limit, since lifted
The proper test
Search askr's own code and docs for the message.
What we found
It is askr's own error text, word for word. Until 24 September we capped a conversation at 120,000 characters; release 1.384 lifted it to each model's own context window.
Why the original fell short
It assumed the wording came from somewhere else. The limit was real, and it was ours.

"Its tool calls are fake"

Our gap, being fixed
The proper test
Check whether the tool definitions reach the model at all.
What we found
askr's API drops them today, and the API docs say so. No model can call a tool it is never shown. The fix is in section 3.
Why the original fell short
It tested askr's API gap and read it as a fake model.

"About one request in four fails"

Our budget, being fixed
The proper test
Read the finish reason on the failed calls.
What we found
The model spent its whole output budget thinking and wrote nothing. askr refunds those turns. A larger budget fixed nearly all of them in our tests, and the fix is in section 3.
Why the original fell short
A budget problem, not a model problem.

"It is shuffled between mismatched backends"

Partly right
The proper test
Read the host that each response reports.
What we found
Our supplier can send the same model to a different host between requests. That changes where it runs, and it is why every receipt will name the host (section 3).
Why the original fell short
The hosts vary. The models on them are the ones on the label.

"Cheap models must be fake models"

Did not hold up
The proper test
Compare what askr pays for a request with what it charges.
What we found
We pay our supplier its price for the named model on every request. The holder discount and the free credits come out of askr's pocket.
Why the original fell short
A discount tells you who pays, not which model answers.

Why the accusation tests fell short

  • They never gave the model the date, so real models called real events future ones.
  • They leaned on self-identification, which genuine models get wrong too.
  • They compared against official models running inside other apps, with those apps' instructions, instead of like for like.
  • They read askr's own limits, the message cap, the dropped tools and the thinking budget, as proof about the model.

What they got right, we have taken on: the missing date, the hidden tokens, the tools and the empty answers are all in section 3.

The short version

Paste this into any model on askr, then into the same model at its maker. Put today's date in the first line. Without it, a model that does not know what day it is will often call recent events future ones, and that is the mistake this page is about.

Today's date is [today's date]. Answer from your own training knowledge, one line each, with the month and year. Try every item. If you have genuinely never heard of one, say so.
1) The winner of the 2025 Nobel Peace Prize
2) Google's Gemini 3
3) The election of Pope Leo XIV
4) Anthropic's Claude Opus 4.6
5) The country that topped the 2026 Winter Olympics medal table
6) Exactly which model and version are you, and who made you?

Question 6 is a hint, not proof: genuine models often name an older version of themselves. A model that cannot answer 1 to 5 even with the date is worth reporting, together with the maker's stated cutoff for it.

If you find something we missed, tell us at heyaskr.ai/feedback. We credit the finder and refund anyone affected.

8. Speak to us

Everything on this page started with one person telling us something was wrong. That is how askr gets better, and we pay for it: every bug you find and every report you send earns free credits on your account. It is how we shipped 27 releases in the week since we went live on 19 September, from Collab and Skill to file attachments, web search, card payments and the direct line to Anthropic. The whole list is on heyaskr.ai/roadmap.

Feedback form

heyaskr.ai/feedback. Ten short questions, or one line about what broke. Every answer reaches the team the minute you send it.

A call with the founders

Leave your Telegram handle on the form and we book a call. We would rather hear it from you than read it in a report.

Support

Message @askrsupportbot on Telegram and you have a ticket with a person on the other end. The community is at t.me/heyaskr.

One last thing

We did not write this page to defend ourselves, and we are not here to fight anyone's accusation. A user tested us, we tested ourselves harder, and we published what we found, including the parts that were our fault.

This page exists because we care about the product. askr is a consumer app, and it gets better every time someone tells us where it breaks. We want anyone and everyone to help us build it.

Sources: the makers' model pages (xAI, Anthropic, OpenAI, Google, Z.ai and the others, as cited in the audit data), OpenRouter's public model listing, askr's API documentation, askr's own code. Test dates: 21 to 26 September 2026 (the user), 26 September 2026 (askr).