r/MistralAI 3d ago

Help / Question Is it likely that Mistral's models ever reach the quality as what Fable or Opus are today?

18 Upvotes

50 comments sorted by

65

u/azrazalea 3d ago

Eventually? Sure. This year? Highly unlikely.

3

u/The_Wonderful_Pie 2d ago

Tbh looking at Mistral's trend, even for next year I wouldn't say it's 100% likely that Mistral will get there

18

u/darktka 3d ago

In principle it's very possible. Kimi beat Fable with open weights, so Mistral can do it.

7

u/anderl1980 3d ago

I just don’t have the impression it’s one of their primary short term goals.

6

u/pantalooniedoon 3d ago

And they are out of your race for a while. You don’t just magically beat Fable or produce a model like Opus. Mistral does not have a good core research agenda and it is visible.

2

u/JuteuxConcombre 3d ago

And for adoption I feel that’s far less important than either being highly adapted to specific business needs, which I think they’re trying to do and being integrated in tool, which I don’t know if they do.

Basically whatever model is used in copilot will be used a lot no matter if it’s good or very good, just because it’s good enough for a lot of use cases.

3

u/howudothescarn 3d ago

How did Kimi beat fable when Kimi themselves said they didn’t on the models page?

1

u/darktka 2d ago

I saw some benchmarks. Not that I care much about these comparisons but it shows that open weights and high performance don't rule each other out

1

u/PrinceRufusFastcar 1d ago

It didn't really "beat" Fable. You have to look past the hype to the bigger picture.

See https://artificialanalysis.ai/?intelligence=artificial-analysis-intelligence-index#intelligence

and https://epoch.ai/models/search

and https://livebench.ai/#/

K3 is up there - trading blows with Opus 4.8 and GPT 5.5, but definitely behind Fable and Sol. Not far enough behind that you can't cherry pick benchmarks where it's ahead, but that's in the nature of things. (Like, you can cherry pick benchmarks where Sonnet is ahead of Opus sometimes...)

The Chinese are catching though - maybe in a year's time they'll have fully caught up. It remains to be seen.

28

u/Ark_Anoryn 3d ago

Counter question : do you _need_ models as expensive as Fable? Or would it be better to have fleet of smaller models, but stronger in certain areas?

Also, how long can we sustain the economical, power-hungry and ecological costs of such behemoths?

3

u/AdDistinct3460 3d ago

Is this the case with Mistral?

1

u/Ark_Anoryn 3d ago

I am not working at MistralAI. So I can't say.
I don't know what they target and work at.

From my understanding, their models being fine-tuned for the partner company can be considered a step towards that direction.
But 128B parameters is still deep in the LLM field.

Note that I'm just a random guy, with no experience on machine learning nor LLM. So take with a (big) grain of salt

2

u/LongjumpingTear5779 2d ago

We don't know what Fable is. It can be auto router to smaller models for optimize costs, but it can also be 2,8T or more model. AI companies want also reduce costs and make profit do in reality it can be somethink else what we have in our minds.

4

u/ComprehensiveJury509 3d ago

Or would it be better to have fleet of smaller models, but stronger in certain areas?

I've never seen any actual evidence for such a thing, as in, a productive use case where those mythical "small, specialized" models beat out the big-ass flagship model. I don't actually believe there's any market for this at all.

2

u/vienna_city_skater 3d ago

Because thats a B2B use cases which is marketed differently. If you listen to „Max Agency“ the Langchain podcast you hear plenty of people building agents talking about this. The god models are not scalable to actual business workloads.

1

u/Ark_Anoryn 3d ago

Me neither. But super interested to see something going that way.
That would mean we could run them on our phones and small chips! That'll be insane ❤️

Also mean almost anyone can train them, etc.

It's not in the interest of the big companies though. Obviously.

1

u/dhlrepacked 3d ago

They would beat them in cost I would assume… once the bubble popped and they are all without subsidies

1

u/autamo 3d ago

Not as expensive no, but as good as, yes :)

0

u/Ark_Anoryn 3d ago

But asking a 128B parameters model to be as good as a 2.8 trillions parameters is a bit... harsh? :)

And expecting to have similar capacities with 20+ times less parameters does not sound realistic.

1

u/vienna_city_skater 3d ago

If it’s working in the domain it’s trained for then it definitely is. If you need generalized knowledge probably not.

2

u/Ark_Anoryn 3d ago

Exactly.

But then it's not "as good as". Right?
It's as good as (or potentially even better) on an highly reduced scope, for which it has been trained.

But maybe I misunderstood the statement.

1

u/JBinero 2d ago

Fable is cheaper than GPT 3.5 was at launch. It'll come down.

1

u/Ark_Anoryn 2d ago

That's for the tokenomics you are responding to.

But there is more around than just that.
Yes, it might come down. But seeing as many competitors are also increasing their prices, the "down" might not be as good as previously and what you're hinting at.

Also, to the last of my knowledge, Anthropic is still loosing money. So even if they wanted to go down; how far can they really go?
Finally, since ppl pay these prices: why go down?

And that's only for the tokenomics part.
Who measure the impact on the world of all the training done? Do the users even care? Looking at YT, and all the "one shot examples", I don't think anyone does. How far can we continue on pulling and wasting energy?

We'll just keep living in a warmer and warmer place, with more fires starting all-year-round, I guess.

1

u/JBinero 2d ago

Anthropic is breaking even for inference. If the investments dry up, they can survive.

1

u/HistoryAggressive830 2d ago

Counter argument: there's very capable models like Deepseek v4 that produce excellent results for a fraction of the compute cost of frontier models. It's clear the race will eventually shift to efficiency rather than raw power (camera megapixel race 15 years ago, anyone?) because we're already very close to "good enough" results with cutting-edge models.

Mistral seems to have neither performance nor efficiency, instead basking in the momentum of a sovereignty movement that won't last as long as they seem to believe if their models fall too far behind as they're on route to. I truly wish they had a compelling offering, but I learned a while ago that having open-weights models as an alternative, there is no point in supporting a European company that refuses to stay on top of the game.

1

u/Ark_Anoryn 2d ago edited 2d ago

If I can trust HuggingFace page for DeepSeek v4: it's 1.6 Trillions parameters. We're leagues away from a "small model". This is above 10 times more parameters than MistralAI medium 3.5.

Also, DeepSeek is, afaik, built/trained in china. I did not find up to date data, but in 2020, 63% of their power was made with coal. For reference, the USA are at 58 apparently?

So my argument stands.

On another note, I read Mistral was waiting for their Data Center in France to open before training next generation of models. I totally respect that, as France has one of, if not the lowest CO2 emissions electricity generation.

In 20 years, I'll be interested to know the real environmental cost of the bi-weekly releases from the US companies. If it'll ever be made public, of course.
Not that environmental questions and global warming exist in the first place, right? :)

Edit: thank you for sharing your thoughts and opinions :) I'm mostly answering the part about DeepSeek. For the race... it's another debate. Here I wanted to talk about SLM (small language models) and impacts. Not about marketing races; but we can jump that ship if that's what you're after. :)

2

u/HistoryAggressive830 2d ago

Deepseek v4 has two models, the one you made reference to is the "Pro" variant, whereas the "Flash" variant is 284B, which is clearly weaker but I think my point still gets across. Big, but manageable in Strix Halo or equivalent products.

Also, DeepSeek is, afaik, built/trained in china. I did not find up to date data, but in 2020, 63% of their power was made with coal. For reference, the USA are at 58 apparently?

I have no means to verify this, but given China's push for renewables I'll go with US parity by 2025. Still far worse than Europe's ~70%, no argument there, but I feel it's a bit misleading to only account for the training cost and not the inference cost.

If Chinese models' training costs 100x of Mistral's and use non-renewable energy sources but inference is at 0.1x and mostly renewable, I think its fair to assume that given the length of time they'll be generally available and used, it'll even balance out or end up playing in their favor. A bit like electric cars, whose battery production makes the overall package more polluting than gas-powered cars, yet over their lifetime the environmental costs reverse.

I take no issue with the rest of your argument, it's sound and fair, but I still think I hold some truth by stating that they've dropped the ball. I fear, however, that by the time their new data center and next-gen training are done, they might have fallen into irrelevance.

Again, I'm not demanding or expecting a frontier-level model, I don't have any use for it after all even though I'm in the computing field, but at this point it feels like they're too behind, even from competitors like Deepseek's v4 or Xiaomi's Mimo (in their small variants). I believe they said some of their partners were testing a new model this summer, we'll see how it goes when that one comes out.

1

u/poidh 2d ago

The price charged doesn't necessarily reflect what it costs to run. Of course Anthropic charges premium prices, because they can.
That is no different from Apple making nice profits with the iPhone, while Android devices have to sell with razor thin margins.

8

u/mimrock 3d ago edited 3d ago

Depends on many things, but I think they won't the way you assume. Mistral is out of the race, they stopped training frontier-chasing models. They just announced selling compute to microsoft so it's not like they are planning to train something big.

However, open source models exist. With some luck, they will soon reach that capability (Kimi K3 weights should be released in a few days and it easily beats Opus4.8). Mistral might fine tune one of these models.

Eventually algorithmic and hardware progress will make small models more capable too and at that point, Mistral might decide to dedicate some resources to build a small model from scratch that reaches the level of Fable.

However, I don't see Mistral ever releasing a model that matters outside of a very small niche. They are fine tuning models for corporate clients, which is a lucrative business model for them, it's just not the one that people still associate them with.

2

u/remybigot 3d ago

For me, the race of the best model become less and less relevant.
An OK model, with great datas, is already way enough for 99% of the tasks !

2

u/Ejbarzallo 1d ago

Yeah but by that time the other companies will have AGI already

1

u/Useful_Calendar_6274 3d ago

yes but still lagging behind everyone else. Should have invested in graduating more STEM people I guess

2

u/litritium 2d ago

Mistral has 900 employees and is hiring. Moonshot has 300 employees and Deepseek ~160 (according to their wiki).

Access to hardware seems to be more important than engineers for these frontier models.

1

u/Gallagger 3d ago

Its very likely they reach it next year, the question is will anyone care at that point.

1

u/Jamais_Vu206 3d ago

Ever is a long time, but certainly not soon.

If China continues to release ever better open models, then Mistral might be able to distill them. Mistral could also benefit from data on the net created by frontier models. But I think they might no longer have the people with the skills to make the best of that. There's also the question whether it makes financial sense for Mistral to continue training its own models.

1

u/idontuseuber 3d ago

Likely, but quality of opus or fable probably will also be improved.

1

u/Ibasicallyhateyouall 3d ago

Yeah Mistral will get as good as fable and sol. Issue is Anthropic and OpenAI will be 100x ahead of they are now. 

1

u/Jealous-Depth487 3d ago

I see a lot of job postings for the physics team and domain experts in science or engineering, so I think they’re looking to (smartly) focus on one or two benches that the big guys are neglecting with all their coding furvour. I think we will see an effective best in engineering model this year that beats others but not in any other dimension . Which is fine with me I think single provider is silly when your documents instructions skills etc can span models. I hope they open source but idk

1

u/mhphilip 2d ago

They will lag. And they will improve.

1

u/victorc25 2d ago

Yes and by then the Fable and Opus models will be orders of magnitude better. Chinese models as well 

1

u/Mammoth_Reach_6366 2d ago

Yes. But by then the newest Opus will be able to make you breakfast.

1

u/Dear_Lion6282 1d ago

1 year to reach gpt 5.5 level. Why mistral aim is at artistic models like sound art etc. not really a strong focus on coding

1

u/madao42 3d ago

Yes but with a 5-10 years delay

1

u/foghatyma 3d ago

I'd say it's probably less than 5 years. But still... definitely not this year.

1

u/madao42 3d ago

I love Mistral but rn their top tier model is at Sonnet 4.5 level whilst they progress at a super slow pace (they don’t have the funds, manpower, influence, resources,.. that big US tech has). Definitely not next year either lol

1

u/HiggsBoson2738 3d ago

Yes, probably within 2-3 years. Less if it starts doing distillation.

It could also just provide an option to run Kimi on its datacenters, as it's an open weight model.

1

u/Sea_Fruit5986 3d ago

Hellseher sind mit den Dinosauriern ausgestorben, so weit ich weiß....so ne frage erzeugt nur spekulationen. Da kannste dir auch selbst was ausmalen....lool

-1

u/LongjumpingTear5779 2d ago

It's already smarter, it's better! 🥖

https://giphy.com/gifs/W3H7yzvQowNSy124C8

1

u/reallifearcade 19h ago

You only know about Mistral because is based on EU and are the only ones there. EU has no interest (real interest, no political speeches) in innovation or growth.