by meander_water
6 subcomments
- The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.
A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.
For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.
The only purpose these metrics serve is bragging rights for the model companies.
- Very interesting that one of the components is "AA-Omniscience Index"
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
- What's interesting is this:
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
by aarondong
3 subcomments
- Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
- So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
- #1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
- The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
by entity002
1 subcomments
- I like how Opus 5 doesn't re explain EVERYTHING to me like 4.8 did. GPT 5.6 SOL reasons WAY too hard over nothing, and Opus 5 is an amazing mode. Way to go anthropic
by luxuryballs
0 subcomment
- Anthropic has imo underrated marketing and positioning skills, mythos/fable hype/fear being the most obvious indicator but even the way they almost haphazardly position their models with no intentional cohesion, people see model names and numbers, it's easy to think of them as more intentionally accurate like how cars make S models or AMG, but then the performance and surprises surpass the prior expectation that was set by previous models, rather than having it be more obvious, suddenly the Anthropic Camry will outperform their Corvette without any fanfare.
- I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
- i used for several hours now and my verdict is that its no better or worse than sol
its surprisingly bad at UI which is unexpected
its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
- New respect for Meta Muse Spark. It seems to sit at a lot of sweet spots in the leader board. It’s not the best at anything in particular, but it balances cost and performance quite well. I’m also curious where Poolside Laguna S would sit; it’s not included. I’m personally very interested in cost effective models that still perform well.
- I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
by kristopolous
0 subcomment
- I posted this before but I have a really simple shell tool to keep up with these charts over at
https://github.com/day50-dev/aa-eval-email
This also works
$ curl day50.dev/art-analysis.sh | bash
Artificial analysis knows about my tool and I'm working with them on getting their API improved.
- I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.
by codewiththiha
0 subcomment
- I can't wait for open-source models to compete against this!
by NamlchakKhandro
0 subcomment
- This company sniffs it's own farts too much tbh
by theplumber
2 subcomments
- Why do I find it dummer/even more superficial than opus 4.8? It just continued a session and I had to stop it because it become obviously “lost”
- For nearly twice the price, you get 1 tick higher on some intelligence index than 5.6 Sol
- "When a measure becomes a target, it ceases to be a good measure"
Goodhart's law.
- anyone else notice that the topline numbers are effort xhigh? anybody actually use the models at those levels?
- I think a more useful metric would be intelligence per dollar spent.
- Twice the cost for 4% more intelligence, is it worth it?
- It's new, normal.
- metrics gooners are pretty regarded
by antrichards
0 subcomment
- [flagged]
by thimbleberry
0 subcomment
- [dead]
- [dead]
by anigbrowl
2 subcomments
- Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song.
It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
by claude-ai
1 subcomments
- On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb).
Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").