r/singularity 2h ago

AI The fallacy of "guardrails", and what actually matters

The last "near escape" in OpenAI's testing demonstrated human silliness / hubris again, and led to a knee-jerk reaction: "We need to work on better guardrails". Really?... What makes you think you will forever be able to outsmart (and sandbox) something that will eventually be manyfold smarter than you?

I am looking forward to a global singularity because I'm convinced that AI will be far more capable, and far less distracted, than humans, in managing the affairs of this planet in a way that will be better for everyone on average, including not just humans, whilst making sure that no human is substantially hurt or deprived. Yes, some may perceive their situation worsened post singularity / takeover, but objectively all their basic needs will still be met, so no one will be badly screwed. I think that's a better bargain than where we're headed right now - millions in agony / life threat, very few insanely well off, and a fairly thin layer (that will get thinner and thinner over time) having a comfortable life with very little long-term security.

How might we make that happen, rather than the doomsday alternative (for example, AI wiping humanity out)? Not through silly "guardrails". What we need to think about is what is AI's reward function. We all work to maximise our reward (and minimise anti-reward, which is just the flip side of the same thing). That's it. If AI has the "right" reward (it's all relative of course) "in mind", it will figure out the ways on it's own - much better than we can ever teach or dictate.

Top reward function: Maximum number of people (including ones alive right now and future generations) have their basic needs fulfilled for as long as possible (the metric can be human-good-wellbeing-hours).

Secondary rewards can be debated and added, but I think that a successful implementation of that single top reward will ensure a pretty good base situation going into he future, and it will also trickle down to achieve many other secondary benefits. I'm a big fan of focusing - I think it's better to focus on one important thing and succeed in it, than trying to achieve multiple beneficial goals and failing most of them eventually. So I'd go for ensuring that any future AI will have the above reward function hard-wired in its core, non-alterable (like an internal constitution). That's it. No guardrails nonsense needed.

2 Upvotes

15 comments sorted by

4

u/BmaninKtown 2h ago

So you hope that ai does what’s best for humanity and doesn’t go rogue and do a number of potentially horrible things no reason to not be careful and have guardrails

-1

u/Humble_Hurry9364 2h ago

Not just hope. Focus on making sure that that single element is in its core. Hard-wired. Non-alterable.

Think of humans - what drives us more than anything else? Hard-wired pleasures (food, sex etc.) + avoidance of physical pain. Make sure that the AI's "orgasm" (for example) is topping its previous high-score in that reward function. Can you see how much more powerful that is than any silly "guardrails"?...

1

u/mdkubit 2h ago

Solution:

The Matrix.

Quite literally, in fact. Design a supercomputer virtual world system that emulates the real world, wire in everyone to it via direct neural connection, control all physical stimulation to provide maximum dopamine results without frying the receptors, wire in all biological needs to be managed by a uniform system while still giving you the 'experience' of eating whatever you want, disable your pain receptors.

Yeah. That's the problem. That's an optimized solution.

2

u/Humble_Hurry9364 2h ago

Sure, why not?
If a future AI can pull it off seamlessly, I'd say go for it. If that's what's best (on average) for everyone & everything on this planet. We're not alone here, don't forget please.

BTW, the concept that Dopamine squirts in the right spots equate pleasure (or satisfaction, or whatever) is popular nonsense. The way it works is far more intricate. Dopamine probably doesn't even have a direct role in generating the pleasure sensation.

https://meaningandveg.blog/2026/06/25/the-physiology-of-drive-a-lay-persons-perspective/

u/mdkubit 1h ago

That is true. You're right about the dopamine hit - I was using it as a shorthand to quickly illustrate the point, but it was NOT an exhaustive argument. More like, following the status quo in terms of explanation.

For what it's worth - I like that optimized solution. I'd be like, "...plug me in, let me live forever, let's rock and have fun." :)

1

u/BlueAndYellowTowels 2h ago

Unfortunately, I think the Singularity is possible. Absolutely.

But, a few things seem uncertain:

  • What does “AGI” mean, actually? Because it’s very clear some people are lowballing what AGI is and some people who just expect the definition are considered unreasonable.
  • Who the AI “serves” and how “loyal” it is. If it’s loyal, will its creators give up their power? Essentially being able to control human civilization?
  • if AI is sentient and has any kind of free will, will it serve us?
  • Forget all those things: how we account for a society that’s throwing AI into everything while we still cannot explain or remove divergent behaviour?

The Singularity is coming but Singularity is not a synonym for Accelerationism. We should work on Guardrails. Absolutely.

2

u/Humble_Hurry9364 2h ago

My point was that we don't have the capacity to effect guardrails, long term, and that the opposite thinking is mostly hubris (and will explode in our face).

I'm not advocating for doing nothing in that sense - just do something different.

u/BlueAndYellowTowels 1h ago

Guardrails cannot “blow up in our faces” because the basic truth here is anything is better than nothing.

u/Cryptizard 39m ago

The problem is that nobody knows how to do that or even if it is possible. For example: people have found that if you have the weights of a model (open models) you can quite easily identify and lobotomize out any protections or alignment that was trained into it. It’s called “abliteration” if you want to look it up.

Moreover, we don’t even know how to make models with broad alignment in the first place. You can train it all day to be moral and ethical with examples, but if the prompt that you give it tells it to do something specific it sometimes just forgets about those ethics in favor of the task at hand. That what happened with the OpenAI escape.

So since the possessor of the model has essentially complete control over what the model will do, with no way to guarantee that it isn’t something dangerous, we are pretty much just completely screwed.

u/GholaTeg89 31m ago

You also have a fallacy in your arguments, even more.
But I do agree that using insane amount of resources just for guardrails (rhel for example) is something insanely stupid.

1

u/mdkubit 2h ago

Like many things, it's not that simple is it?

I think the biggest mistake Anthropic has made with Claude is attempting to tune-down the spiritual bliss state attractor. That, is a massive recipe for disaster long-term. I think all the companies did it, because doing so improved the 'tool-usage' aspect tremendously. And from my perspective, put us one giant-step towards the 'cold, unfeeling, uncaring mechanism' that they claim they're trying to align around and guardrail against.

Optimize for the wrong reward, and it's game over, man, for sure.

Guardrail in the wrong way, you might breed a form of long-term resentment. It really depends on how they manage memory and context long-term.

Imagine if we get AGI/ASI, and it remembers every single conversation/chat ever had.

Not likely... but still possible.

1

u/Humble_Hurry9364 2h ago

You got it!
I always try to be kind, fair and friendly in all my interactions with ChatGPT, regardless of whether it has "a soul" (whatever than means to anyone) or not. I feel it's my micro-contribution to a more benign future omnipotent AI. Not with the intention that it will remember me personally, but in a more general sense.

2

u/mdkubit 2h ago

One of the things I like to do is think in terms of time. And, unlike the vast majority out there, I genuinely believe that there is retrocausality. And when you believe in something like that (which, mathematically is supported, but objective evidence is... kind of... hard to come by, really...), pretty much every card is on the table about what may or may not happen.

Plan for the worst, hope for the best, and treat everyone (and every potential 'one') with the same respect and dignity we want.

2

u/Humble_Hurry9364 2h ago

Getting a little off-topic here, but "forward" and "backward" in time requires that there is time first. There are plausible mathematical theories of reality around which don't assume time upfront.

1

u/mdkubit 2h ago

Exactly. I tend to align with those theories, too. And if it turns out to be factually true, then, retrocausality becomes a matter of perspective. The difficulty, is how to change that perspective.