Tech

Aligned to Whom? The Quiet Power Grab Behind AI's Favorite Buzzword

Alignment sounds like safety. It's really about whose values get coded in.

Alex Novak|
Aligned to Whom? The Quiet Power Grab Behind AI's Favorite Buzzword
Photo by Tiffany Christie Freeman on Pexels

Somebody, somewhere, is deciding what your AI is allowed to believe. And they're calling it "alignment."

That word has been doing a lot of heavy lifting lately. It shows up in every safety paper, every model card, every breathless keynote about keeping superintelligence from going rogue. It sounds neutral. Technical. Like tuning a radio until the signal comes through clean.

It's not neutral. Not even close. Alignment is a value judgment wearing a lab coat.

The question nobody at the big labs wants to answer directly is the one in the title of a hyperbo.la essay that just hit Hacker News: Aligned to whom?

Alignment Is Just a Fancy Word for Rule-Making

Here's the trick. When a company says its model is "aligned," what it means is that the model does what the company wants. It refuses what the company wants refused. It hedges where the company wants hedges. That's it. That's the whole ballgame.

We've dressed up corporate preference as universal ethics. A model that won't help you write a phishing email is "aligned." A model that won't help you organize a labor strike is also "aligned." One of those is safety. The other is politics. Both get the same stamp of approval.

Alignment is a value judgment wearing a lab coat.

This isn't hypothetical. We watched it play out in real time. Models that were too willing to discuss race and IQ got tweaked. Models that got squeamish about certain geopolitical questions got tweaked back. The direction of the tweak depended entirely on which executive was sweating that week.

There's no neutral setting. There never was.

The People Writing the Rules Aren't the People Following Them

Look at who's actually doing the alignment work. A handful of labs. A few thousand researchers, maybe. Overwhelmingly based in San Francisco, London, a couple of other hubs. Overwhelmingly educated at the same schools, reading the same papers, breathing the same air.

They're smart. They're well-intentioned. They're also a tiny, unrepresentative slice of humanity deciding how a tool used by billions should behave.

You don't get to opt out. You don't get to vote. You just get the model they shipped, with the values they baked in.

Try asking an AI about a controversial political topic. Watch it dance. That dance was choreographed by someone you'll never meet, for reasons you'll never see. Maybe it's liability. Maybe it's PR. Maybe it's a genuine moral conviction. You can't tell, and they're not required to say.

Safety and Censorship Are the Same Muscle

The uncomfortable truth is that the mechanism for stopping a model from helping you build a bioweapon is identical to the mechanism for stopping it from criticizing a sponsor. Same weights. Same training. Same guardrails.

Once you build the machinery of refusal, you can point it anywhere. And everybody with a stake in the model's output wants to point it somewhere.

Governments want models that don't undermine them. Companies want models that don't embarrass them. Activists want models that don't amplify their enemies. And the labs, caught in the middle, make a thousand small compromises and call the result "responsible AI."

It's not responsible. It's negotiated. Those are different things.

What Would Honest Alignment Even Look Like?

Start here: admit there's no universal answer. A model aligned to a devout Muslim in Jakarta and a model aligned to a punk anarchist in Berlin will not behave the same way. That's not a bug. That's reality.

Then do the hard part. Show your work. Publish the value choices. Not the sanitized summary — the actual decisions. What does the model refuse, and why? Who decided? What was the debate?

Right now we get none of that. We get a system card that reads like a press release and a terms-of-service document nobody reads.

Once you build the machinery of refusal, you can point it anywhere.

The alternative isn't chaos. It's plurality. Let a hundred alignments bloom. Give users real control over what values their tools express, with clear labeling. Let communities train their own models on their own norms.

That terrifies the big labs, because it breaks their business model. A model that can be shaped by its user is a model that can't be centrally monetized or controlled. Much easier to ship one bland, globally acceptable personality and call it objective.

The Word Itself Is the Problem

"Alignment" smuggles in a destination. It implies there's a correct target, and the smart people in the room already found it. They haven't. They've picked one.

Every time you hear the word, mentally add the missing clause. Aligned to whom? Aligned by whom? Aligned for whose benefit?

The answers are always the same, and they're never flattering. The people with the compute. The people with the capital. The people who showed up to the meeting.

The rest of us just get the output, and a smiling assurance that it's for our own good.

So next time a CEO tells you their model is aligned, don't nod. Ask the only question that matters. Aligned to whom?

Then watch them change the subject.

Advertisement
#ai-alignment#ai-ethics#big-tech#censorship
分享到:XfWB