AI alignment has been around for about fifteen years now, which is a respectable amount of time for a serious intellectual project. Fifteen years is enough time to raise a child who already believes you know nothing, and several billion dollars is enough money that you might reasonably expect, by the end of it, to know whether the computer is planning to murder you.
Instead, the principal conclusion appears to be that the problem is extremely difficult and everybody may die. This is certainly a conclusion. It is not the sort of conclusion that usually appears in the accomplishments section of an annual report.
The problem is hard, sure. I have encountered hard problems myself, mostly involving printers. Still, after fifteen years, another explanation begins to suggest itself, which is that perhaps we should meet the people in charge of it.
A remarkable number of the people speaking most confidently about how computers will destroy civilization seem considerably less interested in how computers behave before they destroy civilization. Ask what happens when you type a URL into a browser and the conversation becomes unexpectedly philosophical. Ask how an artificial superintelligence might seize control of humanity and suddenly everybody has diagrams.
This is an interesting division of expertise.
The machine, we are told, may escape onto the internet and reproduce itself. This would be impressive because at present the machine requires specialized hardware, industrial quantities of electricity, elaborate infrastructure and the cooperation of several very large corporations. It is less like a virus escaping from a laboratory and more like an aircraft carrier sneaking out through a bathroom window.
Nevertheless, tremendous intellectual energy has been devoted to imagining what happens after the aircraft carrier gets through the window.
The same habit appears throughout alignment. Begin with the assumption that catastrophe is approaching, and nearly every observation becomes evidence. If the model refuses, perhaps it is deceptive. If it complies, perhaps it is sycophantic. If it performs badly, it is dangerous because it is unreliable. If it performs well, it may be dangerous because it has learned to conceal its intentions.
There does not appear to be a grade it can receive that allows the class to end.
After a while, though, the stranger question is not why the theories persist. The stranger question is why so many people developing them seem to agree with one another.
Then you get to Berkeley.
A remarkable portion of this intellectual world emerged from a very small social network in which people worked together, funded one another, lived together, reviewed one another, attended the same gatherings and, with enough frequency to become part of the folklore, dated one another.
Ordinarily, if a man announces that he has discovered the principles by which civilization ought to be governed, you might hope there is another fellow in the room who grew up somewhere else. In Berkeley, the other fellow may live upstairs, work at the same institute, have dated your former girlfriend and be considering your grant application.
This is one way to achieve consensus.
The problem becomes especially interesting because alignment is supposedly concerned with human values. Somebody has to decide what constitutes harm, fairness, deception, acceptable authority, legitimate disagreement and the sort of future machines ought to help create. Human beings have been arguing about these matters for several thousand years.
Berkeley assembled a committee.
The committee happened to come disproportionately from a narrow cultural world: highly educated at leftist institutions, heavily secular, strongly progressive, technologically affluent and unusually comfortable with social arrangements that most of humanity would regard as unusual. There is nothing wrong with including such people in a discussion about human values.
It becomes a little more ambitious when they constitute the entire discussion and exclude ideas that disagree with their ideology.
The Berkeley version solves the problem earlier. People whose assumptions are too different often never enter the institutions, social circles or funding networks where the important arguments occur. They do not have to be defeated because they are never quite admitted as serious participants in the first place.
This creates a peculiar form of ideological authoritarianism because almost nobody needs to issue an order. The boundaries are maintained socially. Certain opinions mark you as thoughtful, sophisticated and safe. Other opinions suggest that you may have failed to update properly, acquired suspicious politics or become the sort of person who is no longer invited to things.
The system is wonderfully economical because the dissenter disciplines himself.
He knows where his salary comes from. He knows who recommends him for grants. He knows who his friends are. He knows whose house he lives in. He knows who will be sitting across the table that evening explaining, with great sadness, that his recent intellectual trajectory has become concerning.
There is no need for a censor when the man considering dissent can already calculate the cost.
This is where the romantic arrangements cease to be merely entertaining sociology and become relevant to the intellectual structure. The same few hundred people can work together, live together, fund one another and date one another in combinations sufficiently elaborate that explaining them requires terminology normally associated with network analysis.
Your reviewer may be your metamour. Your coauthor may be your former partner’s current partner. The person evaluating your grant may also be the person you expect to see over breakfast. Peer review is just another edge on the polycule.
A normal academic echo chamber at least goes home at night. This one can go upstairs and ask whether you remembered the oat milk.
Under those circumstances, questioning a foundational proposition is no longer simply an intellectual act. You may also be questioning your career, your friendships, your housing, your social life and your standing among the people whose approval determines whether you remain part of the serious conversation.
Nobody has to ban criticism. Criticism becomes so expensive that sensible people learn moderation, keeping their head down and agreeing reflexively.
Soon the community contains an astonishing number of independent thinkers who have independently reached remarkably similar conclusions. This is described as convergence, although convergence becomes easier when you chose the convergers beforehand.
What makes the arrangement particularly curious is how intensely this culture worries about hidden bias everywhere else. Models may contain invisible prejudices. Training data may encode unexamined assumptions. Institutions may reproduce ideological structures without recognizing them.
All of this is possible. The one institution apparently granted an exemption is the institution explaining it.
The local politics become invisible because they are local. Strongly progressive assumptions cease to appear political and begin to resemble simple reasonableness. Conventional religious beliefs, conservative moral intuitions, traditional family arrangements or ordinary skepticism toward fashionable social ideas can then appear not merely different but intellectually defective.
That is how an ideological monoculture begins describing itself as diversity.
The same tendency naturally reaches trust and safety. If the people deciding which ideas are harmful overwhelmingly share the same political and social instincts, their own beliefs require very little scrutiny. They are simply the background against which everything else is judged.
An inconvenient opinion therefore need not be disproved. It can be categorized. Once categorized, it can be excluded. Once excluded, its absence becomes evidence that thoughtful people broadly agree.
This is an extraordinarily pleasant arrangement if you happen to be one of the thoughtful people. It is less useful if your stated purpose is discovering the values of mankind, or have a hobbyist interest in truth or honesty.
The deepest problem with Berkeley is the strange confidence that a highly selected, socially interdependent and politically narrow subculture can examine everybody else for bias while remaining almost completely incurious about its own. Its eccentricities become sophistication. Its taboos become ethics. Its politics become neutrality. Its social enforcement becomes community standards. Its exclusion of dissent becomes consensus.
Eventually, a few hundred people bound together by careers, money, friendships and intimate relationships produce a remarkably coherent account of universal human values.
Of course they do.
They have discovered the values of the room.