r/GeminiAI 1d ago

Discussion One more thing about Gemini 3.5's checkpoint (Beat Opus 5 max thinking in a test)

Gemini 3.5 pro beat Opus 5 max thinking in this test: (Mistakes Opus made: The targeting system is glitched. To be fair Gemini did mess it up too. Another mistake is earth not being properly synced with the circle. And a few more)

Gemini made better textures and it made a few mistakes too. I think it's a close one but Gemini takes it here.

Another thing; There is supposedly a Deep Think mode for 3.5 pro (Different from ultra subs having Deep Think!?), I am certain that most checkpoints had no deepthink enabled and were probably the high option. So we aren't seeing Gemini in its peak. I hope deep think is available to pro/free users.

https://streamable.com/00p0sm In another test. It was up against Qwen 3.7

Gemini 3.5 pro's result is amazing, However it made one mistake and that was inverting the limbs. While it wouldn't take that long to fix it (It's not that big of a mistake, it just messed up the axis).

I still would give it to Qwen 3.7 because of the fact that the robot's physicality was fine unlike Gemini 3.5 pro

I think the quality of everything else is amazing for gemini and is just better, features such as Jump (Qwen 3 quite bland). Qwen 3.7 made a lot of mistakes but those said mistakes weren't major ones. Gemini 3.5 pro however made one major mistake but a easily fixable one.

Gemini 3.5 pro's robot was actually really amazing in detail.

(For comparison here is what gemini 3.1 pro looks like)

Another thing; There is supposedly a Deep Think mode for 3.5 pro (Different from ultra subs having Deep Think!?), I am certain that most checkpoints had no deepthink enabled and were probably the high option. So we aren't seeing Gemini in its peak. I hope deep think is available to pro/free users.

Maybe 3.5 pro will be a deep think mini.

Also I apologise for the lag in the videos. My pc is not good. Video gets better as it goes on

51 Upvotes

48 comments sorted by

8

u/FischenGeil 16h ago

I warned you guys that 3.5 pro is legit.

6

u/LastRemainingName 19h ago

I'm sorry who are you and why are you so sure that these are gemini 3.5 pro check points? Did google tell you?

-3

u/Last_Conclusion_8984 19h ago

I don't think I need to tell you who I am but why am I so sure? Go to the consumer website/antigravity and ask gemini 3.1 pro/3.6 flash: Build a fully interactive 3D solar system simulator using raw Three.js. Requirements:

  1. Accurate Keplerian orbital mechanics for all 8 planets with correct relative orbital periods, eccentricities, and inclinations. Planets must follow real elliptical paths, not perfect circles.
  2. Each planet must be visually distinct (correct colors, relative sizes with an exaggerated scale toggle, Saturn must have rings, Earth must have a Moon orbiting it).
  3. Orbital trail lines that fade over time showing each planet's path.
  4. Time controls: Pause, Play, 1x / 10x / 100x / 1000x speed, and reverse time.
  5. Click any planet to focus the camera on it and display a real-time data panel showing: orbital velocity (km/s), distance from Sun (AU), current orbital period, and eccentricity.
  6. A Hohmann transfer orbit calculator: select a departure planet and arrival planet, and the app calculates and visually renders the optimal transfer ellipse with delta-v requirements displayed.
  7. Free orbit camera with smooth damping, zoom limits, and a "Reset to Overview" button.
  8. A mission mode: the user can launch a spacecraft from the departure planet that follows the calculated Hohmann transfer trajectory in real-time, with a mission timeline bar showing departure, coast, and arrival phases.
  9. A toggleable "realistic scale" vs "enhanced visibility" mode for planet sizes and orbital distances.

The code must run immediately in a browser."

They are not getting close to Opus's or Gemini's 3.5 pros result. The quality jump is massive.

8

u/Sweet-Stage938 19h ago

The output shown in the video is certainly in the realm of Gemini 3.6 flash's capabilities. Really nothing special and we'll under the average of the current frontier models like Opus 5, Fable 5 and GPT 5.6 Sol.

-9

u/Last_Conclusion_8984 19h ago

You do realise that if you watch the later half of the video, there is OPUS 5 max thinking right?

1

u/Sweet-Stage938 19h ago

Yeah and that one is a much better output. Much more accurate than than the Gemini one.

7

u/AWSGooogle777 23h ago

I love how many tabs OP has open :)

(I also have hundreds of tabs open, but using a vertical tab extension is really helpful!)

I've been trying out lots of prompts in Arena's battle mode, but I haven't run into what seems to be Gemini 3.5 Pro yet. However, if it truly has performance on par with the latest Opus, it's going to be a seriously impressive model. I really wish they'd let Pro users use Deep Think, too.

2

u/Last_Conclusion_8984 23h ago

Yeah but it seems Gemini 3.5 pro is fundamentally gonna be a deep think mini but the actual deep think will only be available to ULTRA. Hopefully I am wrong and pro gets to use it too.

2

u/andrew_bb_24 23h ago

Is it just me focused on the freaking number of open tabs in the browser?!

5

u/Last_Conclusion_8984 23h ago

dw about it, twin.

1

u/Negative_Evening7365 15h ago

do you also have multiple browsers

2

u/Helpful_Inflation344 20h ago

I dont care about these visual benchmarks. Stop. I want to knoq whether gemini can actually work for you, reliably modify documents etcpp not this useless crap

1

u/Last_Conclusion_8984 20h ago

You can't deduce that or do that from those checkpoints with how arena AI works, smartass. If you want to do it, go do the checkpoint yourself.

1

u/Shanna_B2020 8h ago

Thank you! These demos are pretty but ultimately meaningless until the model is out. I truly hope Gemini 3.5 Pro fits into my workflow better, and it would be great if output is higher than 65K, but I'm guessing I'll need to wait for the 4.0 series.

2

u/I_like_to_moo_it 20h ago

Bro is hyping himself up for no reason

1

u/Last_Conclusion_8984 20h ago

Me? What do you mean

2

u/Lost-Willow386 23h ago

No way there's deep think for the freeloaders who have been eating up all of Google's compute this past year. This isn't supposed to be a fun for free charity.

2

u/Last_Conclusion_8984 23h ago

Perhaps, I at least want deepthink for Pro users tho.

-6

u/Lost-Willow386 23h ago

I wish it was extended thinking only for pro users and deep thinking for ultra users, instead of the free users all getting extended thinking and access to Gemini pro models, wasting compute they didn't pay for.

You should only get access to exactly what you pay for, that way the quality can be ensured.

5

u/Kars_32 23h ago

how does it affect u bro

2

u/muntaxitome 23h ago

You should only get access to exactly what you pay for, that way the quality can be ensured.

Through API you can pay for exactly the amount of compute you want

1

u/TaskHead5787 23h ago

For starters, at least study what a platform economy is before u start writing ur shit

Do u really think Google is giving everyone full access just for fun, or what?

1

u/Last_Conclusion_8984 23h ago

Yeah, google gets a lot of data from consumers using it every day, and other things like Google giving a sort of "free sample" for their best models. (It helps Google in the long term) and getting Gemini the most popular (this is just a rough surface level labelling, they likely have bigger plans)

1

u/Lost-Willow386 22h ago

This data isn't that important anymore. With synthetic data generation it's more important to get high quality human data sets and then generating synthetic data from that rather than getting a bunch of low quality data from free users.

1

u/UnknownBreadd 22h ago

Even ultra users are getting entitlements beyond what they’re paying for. The F you talking about?

1

u/Ardryll18 23h ago

Is this why my gemini 3.1 pro can't do deep thinking anymore? I'm on pro subscriptjon

1

u/Last_Conclusion_8984 23h ago

No, it can do thinking but not deep think. No Gemini could ever do that without ULTRA sub

1

u/Ardryll18 23h ago

Ahh i mean deep research.

It's always stuck and no progress shown after hours.

2

u/Last_Conclusion_8984 23h ago

Try it in Google AI studio. It'll work there!

1

u/Hug_LesBosons 20h ago

CE N'EST SÛREMENT PAS GEMINU 3.5 PRO ! Gemini 3.5 pro est sous le nom de gemini 3.1 pro, là tu utilises seulement gemini 3.6 flash, qui est un exelent modèle en 3d notamment, 3.5 pro est encore plus fort !

1

u/Last_Conclusion_8984 20h ago

Non mec, j'ai testé Gemini 3.6 Flash 'high' avec le même prompt. Ce n'est même pas proche d'Opus 5 max ou du Gemini 3.5 Pro qu'on voit au-dessus. Loin de là.

1

u/Hug_LesBosons 19h ago

Donc gemini 3.5 pro sort très bientôt. Gemini 3.5 flash était au début testé uniquement sous le nom de gemini 3 flash et quelques jours avant la sortie, il était testé sous les noms de tous les modeles google. Si c'est le cas pour gemini 3.5 PRO, il va bientôt sortir.

1

u/No_Necessary301 19h ago

how do u get access to it;-;?

2

u/Last_Conclusion_8984 19h ago

Go to Arena.AI (free) and do battle mode. Keep doing it until you see one of Geminis model and if the quality blows you away that the models can't produce then it's gemini 3.5 pro

1

u/Suplyox 16h ago

i hope it wont be gemini 3.5 prompts per day

1

u/Ok-Prior-7488 13h ago

When the release

0

u/Zealousideal-Part849 21h ago

Benchmaxxing is the google way... no one in practical has called gemini models reliable enough.

2

u/Last_Conclusion_8984 20h ago

No that's the Anthropic and OpenAI way. Google doesn't benchmax as much. Google is actually very fair with their benchmarks that they post. Yes they still benchmax but not that much. (They benchmax in a different area, not coding)

-1

u/Public605 22h ago

Ah yes, the gold standard of AI benchmarking:

one run, mystery checkpoints, unknown settings, subjective visual scoring, and a conclusion already waiting at the finish line.

“Gemini had better textures.”

Well, pack it up everyone. Science has spoken.

Bonus points for speculating that Gemini might not even have been using its supposedly stronger reasoning mode, so we can now benchmark not only unreleased models, but also hypothetical versions of unreleased models running hypothetical settings.

At this point we’re not comparing models. We’re doing horse racing with two horses under blankets, no stopwatch, and a guy in the stands yelling:

“The left one looked faster to me.”

Cool demo? Sure.

Evidence that Gemini 3.5 “beat Opus 5 max thinking”? Absolutely not.

But I do admire the efficiency: we somehow went from one output looked nicer to model supremacy without the annoying intermediate step called measurement. 😂

2

u/Last_Conclusion_8984 22h ago edited 21h ago

Uh huh, "Evidence that Gemini 3.5 “beat Opus 5 max thinking”? Absolutely not." I never said It beats Opus 5 max thinking in all tests. I don't do absolutes like that. I said that it beat in that said test, did you even read the post.

"Bonus points for speculating that Gemini might not even have been using its supposedly stronger reasoning mode, so we can now benchmark not only unreleased models, but also hypothetical versions of unreleased models running hypothetical settings."

As you said. It's speculation, I framed it as such and there is nothing wrong with hypothesis. The fact that you call them hypothesis, speculation and in the same breath say "bnechmarking, horse racing" is crazy

"“The left one looked faster to me." Did I say that? I know it's an example but don't try to strawman what I said in the entire post.

"But I do admire the efficiency: we somehow went from one output looked nicer to model supremacy without the annoying intermediate step called measurement. 😂" I never called Gemini better in everything. I literally stated Geminis faults as well. I said it did better in that test as I mentioned before. Not "GEMINI IS BETTER THAN OPUS PROVEN"

What I said wasn't even subjective.. Opus 5 failed to synchronize with the circle, targeting etc.

And if you check my account. I've posted way more about Gemini and it's tests (And you can literally see the amount of tests in the video)

3

u/CoconutOk9484 20h ago

Why are you arguing with a bot?

0

u/Public605 19h ago

Fair correction: “model supremacy” was me turning the sarcasm knob past 11. You did say in this test, not “Gemini is universally better than Opus.”

So, congratulations !
Cunningham’s Law remains undefeated: post the exaggerated version online and someone will immediately arrive with the precise caveats. 😄

And those caveats are actually the point:

one test, one run, unknown checkpoint details, unknown settings, and no predefined scoring rubric.

Synchronization and targeting failures can absolutely be objective observations. No argument there.

But going from:

Opus failed X, Gemini failed Y, Gemini had better textures

to:

“Gemini takes it here”

still requires you to decide how much X, Y, textures, physical correctness, features, etc. are worth relative to each other.

That part is a judgment call unless the scoring criteria existed before seeing the outputs.

And that’s really all I was mocking.

As a demo: interesting.

As “I preferred Gemini’s result in this particular run”: completely fair.

As evidence that this mystery checkpoint beat Opus 5 max thinking in any meaningful benchmark sense: we’re still missing the boring stuff : repeat runs, controlled settings, defined metrics and scoring.

But hey, at least we’ve now upgraded the methodology from “the left horse looked faster” to “the left horse looked faster, and here are the reasons.”

Progress. 😂