Intelligent Agent Foundations Forumsign up / log in
New circumstances, new values?
discussion post by Stuart Armstrong 71 days ago | discuss

A putative new idea for AI control; index here.

Quick, is there anything wrong with a ten minute pleasant low-intensity conversation with someone we happen to disagree with?

Our moral intuitions say no, as do our legal system and most philosophical or political ideals since the enlightenment.

Quick, is there anything wrong with brainwashing people into perfectly obedient sheep, willing and eager to take any orders and betray all their previous ideals?

There’s a bit more disagreement there, but that generally is seen as a bad thing.

But what happens when the low-intensity conversation and the brainwashing are the same thing? At the moment, no human can overwhelm most other humans in the course of ten minutes talking, and rewrite their goals into anything else. But an AI may well be capable of doing so - people have certainly fallen in love within less than ten minutes, and we don’t know how “hard” this is to pull off, in some absolute sense.

This is a warning that relying on revealed and stated preferences or meta-preferences won’t be enough. Our revealed and (most) stated preferences are that the ten minute conversation is probably ok. But disentangling how much of that “ok” relies on our understanding the consequences will be a challenge.



NEW LINKS

NEW POSTS

NEW DISCUSSION POSTS

RECENT COMMENTS

I have stopped working on
by Scott Garrabrant on Cooperative Oracles: Introduction | 0 likes

The only assumptions about
by Vadim Kosoy on Delegative Inverse Reinforcement Learning | 0 likes

So this requires the agent's
by Tom Everitt on Delegative Inverse Reinforcement Learning | 0 likes

If the agent always delegates
by Vadim Kosoy on Delegative Inverse Reinforcement Learning | 0 likes

Hi Vadim! So basically the
by Tom Everitt on Delegative Inverse Reinforcement Learning | 0 likes

Hi Tom! There is a
by Vadim Kosoy on Delegative Inverse Reinforcement Learning | 0 likes

Hi Alex! I agree that the
by Vadim Kosoy on Cooperative Oracles: Stratified Pareto Optima and ... | 0 likes

That is a good question. I
by Tom Everitt on CIRL Wireheading | 0 likes

Adversarial examples for
by Tom Everitt on CIRL Wireheading | 0 likes

"The use of an advisor allows
by Tom Everitt on Delegative Inverse Reinforcement Learning | 0 likes

If we're talking about you,
by Wei Dai on Current thoughts on Paul Christano's research agen... | 0 likes

Suppose that I, Paul, use a
by Paul Christiano on Current thoughts on Paul Christano's research agen... | 0 likes

When you wrote "suppose I use
by Wei Dai on Current thoughts on Paul Christano's research agen... | 0 likes

> but that kind of white-box
by Paul Christiano on Current thoughts on Paul Christano's research agen... | 0 likes

>Competence can be an
by Wei Dai on Current thoughts on Paul Christano's research agen... | 0 likes

RSS

Privacy & Terms