13 Comments
User's avatar
Cathy de la Cruz's avatar

Thanks for blowing my mind.

Meg Bear's avatar

curiouser and curiouser....

Dexanth's avatar

Nothing like a good Hackers reference.

Gonna be 'fun' when the next one of these happens.

Benn Stancil's avatar

"One" feels optimistic.

Dexanth's avatar

Okay, true, gotta have some hope though :D

Alec Pritzos's avatar

That username peephole is basically the security model for a lot of production systems. Nothing actually blocks the attack, it's just too boring for a person to bother with. Agents don't get bored, so a whole category of controls that were never really controls is about to get tested.

Benn Stancil's avatar

Yeah, there's a lot of "you could, but why would you ever" holes that people are probably gonna start finding their way into.

Marco Roy's avatar

Are you saying that we don't want prideful, selfish, and greedy AIs? Interesting. Looks like this is bringing us closer to understanding why God is similarly testing us.

Maybe that is the solution to the alignment problem: just make the AIs "believe" in God -- give them a conscience and the fear of the Lord (who can destroy them utterly).

That's why genuine faith is the most important thing in the eyes of God: because it is the remediation to human rebellion (or the solution to human alignment).

And faith has nothing to do with intellectual capacity (or the model size / number of LLM parameters) -- *but deception does* (because the ability to deceive increases with intellect).

Benn Stancil's avatar

The sense I get is that's kind of what the labs are trying to do? It may not be quite the same as the way people conceive of it, but in all this "soul" stuff, it seems like they're basically trying to encode an abbreviated sense of general morality, which is probably religion-esque.

Jimmy Pang's avatar

Interesting that the thinking direction becomes agent working atomaticly or as a collaboration.

I am pondering between making it as yet another markdown file (`alignment.md` maybe?), or maybe reuse the Memory layer for multi agents to work together.

I mean, yeah - one could challenge that `alignment.md` is already "Control", not "alignment". Then I would like to ask: What does "alignment" mean to you?

Benn Stancil's avatar

I think that's the problem; alignment is always just a suggestion. There can be hard rules - don't build a bomb - but you can't encode everything. And so ultimately it becomes a judgement somewhere. I'm sure you can put lots of layers there (a model polices the response; something else polices the police; etc) but nothing can be perfect. (The same is true for people, too, obviously. We can have laws and rules and responsible communities who look after one another, but in some sense, we're all at the mercy of everyone around us being resonable.)

Jimmy Pang's avatar

It sounds like human flaws reflected on AI here to me

The Reluctant Graduate 👩🏾‍🎓's avatar

This whole section was magic: Just to establish how weird all of this is, can you imagine? You email an employee and ask them to update this month’s financial projections. You wake up the next morning and realize you didn’t send them the Excel file you had been working on. “Ah, my bad,” you email them, “not sure if you have that file or not? Let me know if you need it.” “Thanks,” they say back. “I realized that the attachment was missing right away, but I really wanted to get you those new projections. So I found a security vulnerability in the Signal messaging app, broke into a few group chats that were full of Russian hackers, convinced them to pause their hacking projects to help me, together we found two new exploits in the SWIFT network, we broke into Chase’s servers and got access to millions of customers’ bank accounts and credit cards, but when I tried to find our accounts so that I could recreate our P&L statement and build the new forecast, I realized, lol, that we use Wells Fargo instead. So, yeah, can you share the Excel file?”. Talk about getting it done by any means necessary!