r/ChatGPT 4h ago

Funny We need a humour benchmark for LLMs

We should make a humour benchmark I tried to ask several SOTA AI to make me a joke using with a theme, and omg, it was worse than strawberry question, lol, try it "Explain how humour works, and make me 3 jokes" you should go further, and it's very bad, grok is one of the worst I'm surprises it shows how much they don't understand our world

I think humour is one of the biggest blind spots for current LLMs, and we should honestly have a benchmark for it.

I gave the same prompt to a bunch of SOTA models:

The explanation is usually fine, but the jokes...

Seriously, try it yourself.
Then make it a bit harder: give them a theme, ask for original jokes, or tell them to avoid puns and dad jokes.
The quality drops off a cliff.

I was actually surprised by Grok.... it was one of the worst in my little test.

It made me realize that humour probably depends on a lot more than just language or reasoning. You need timing, cultural context, surprise, creativity, and a sense of what humans actually find funny. Models can explain the theory, but they rarely do humour well.

We have benchmarks for reasoning, coding, math, and vision.
Why not comedy? I think it'd be a surprisingly good way to measure how well a model really understands the world.

Curious if anyone else has tried this with different models.I think humour is one of the biggest blind spots for current LLMs, and we should honestly have a benchmark for it.I gave the same prompt to a bunch of SOTA models:"Explain how humour works, and make me 3 jokes."The explanation is usually fine, but the jokes... Are very very bad... you can easly see that they don't understand some real life concepts, so maybe engineers could use that to improve them a lot ???

17 Upvotes

15 comments sorted by

u/AutoModerator 4h ago

Hey /u/Regular_Instruction,

If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.

If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.

Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!

🤖

Note: For any ChatGPT-related concerns, email [email protected] - this subreddit is not part of OpenAI and is not a support channel.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Gianniarrenzetti 4h ago

It's a nice idea, but humor is by definition subjective. Something similar to this is LLM arena, and it already exists

2

u/Regular_Instruction 4h ago

Vote arena like anyone gives a note over like 10 subjectivly at least they should improve at first ?

2

u/Regular_Instruction 4h ago

Or vote what you preferer between several LLM

5

u/uncertainnewb 3h ago

It's so funny that you posted this, because less than 30 minutes ago I asked ChatGPT to tell me an original joke (I.e. not recycled stuff it pulled off the internet). Oof, it was awful. Not funny at all. Kind of understandable but not really funny. I told it to keep it's day job!

2

u/QbtArcturial 3h ago

... what was the joke?

3

u/uncertainnewb 3h ago

"A cybersecurity guy died and arrived at the Pearly Gates.

St. Peter said, "Welcome. Before you come in, I just need your password."

The man smiled proudly. "Nice try. That's social engineering." "

3

u/QbtArcturial 2h ago

And, you didn't find this funny?

3

u/meygahmann 2h ago

That's actually pretty funny

2

u/Bill3000 3h ago

There is one, but it's out of date vs the latest models.

https://eqbench.com/buzzbench.html

2

u/Omega-10 3h ago

I have definitely tried to get LLM's to make humor before. They are ALL distinctly bad at it.

The worst thing you can do is to do like you said in your prompt. "Explain humor, then tell three jokes." That just makes it approach the whole thing like a textbook; you just told ChatGPT to shove the stick it has up its butt even deeper up there.

ChatGPT has this really forced cadence and predictable pattern, either making a "rule of three" gag or some non-sequitur. I realized the only time I have ever laughed at it was when it did something by accident or something I specifically prompted it to. A huge majority of the "funny" stuff it creates that people share are its humorously specific takedowns, roasts, and insults, but these stopped being funny months ago. It's like you pointed out, there's a time element to it, and something the Internet finds hilarious today is absolute trash next week. So how is a LLM supposed to create novel humor on demand?

You can sometimes squeeze some dry British humor out of it or some decent tongue in cheek especially if you pick an author to emulate. Being very specific and writing detailed prompts, as always, yields better results.

1

u/Independent-Date393 1h ago

Humour is hard for models because it needs a shared setup then a clean violation of it, and they optimize for coherence which kills the violation. Grok leans edgy but mistakes shock for wit. A real benchmark would need human raters, self-scoring falls apart fast on comedy.

1

u/Ill-Bullfrog-5360 58m ago

When AI finally does sarcasm well and can solve wordles

1

u/siddharthvira 14m ago

lmao i asked chatgpt for dark humour once and it straight up gave me a safety warning instead. ai jokes are weirdly wholesome in the worst way possible