Skip to content
Churches are selling their land for big bucks

Y.I.G.B.Y. (Yes In God's Backyard)

Your pension benefits might not follow you home

Why the residency trap forces immigrants to age in the cold

Search

Latest Stories

What does it even mean for AI to be 'aligned'?

As AI apocalypse fears circulate and tech companies tout the dawn of AGI, aligning machines with human interests is more important than ever. If only we could define what 'human interests' even are.

An animation of the sculpture The Thinker by Rodin with data strings coming out of its head

As Silicon Valley races to create 'AI alignment,' maybe we should figure out human alignment first.

Getty Images


Last week’s much-hyped launch of GPT-6 Astra, OpenAI’s newest model, has been a bit overshadowed by headlines about AI doomsday scenarios and the importance of solving “AI alignment.” Just days after the model was announced, OpenAI’s Chief Scientist Jakob Pachuki penned a blog post for the company’s website, writing that “the core problem in AI research is that of alignment – getting the AI to 'try to do the right thing' by human standards.” In Pachuki’s opinion, “no lab has solved alignment.”

Former OpenAI and Anthropic researcher Jacob Coxon appears to agree: He publicly resigned from Anthropic on Tuesday, posting on X that “neither company is acting responsibly” and that “the people building AI earnestly believe that it could kill us all by the end of the decade.” Current Anthropic researcher Evan Hubinger, who could use some 'positive spin' lessons, chimed in: “we really do earnestly believe AI could kill all humans!” (Was the exclamation mark really necessary, Evan?)

AI extinction fears aren’t new, but GPT-6 Astra is being released to the public despite OpenAI admitting that it meets their 'critical cybersecurity threshold' – not a good thing. Its own Preparedness Framework instructs researchers to “halt further development” if a model reaches this threshold, at least until sufficient safety standards have been put in place. Maybe it’s just me, but when it comes to GPT-6 Astra, it doesn’t seem like very much has been halted. (To its credit, OpenAI did pause testing for one of its models – not Astra – after the now-infamous Hugging Face incident… for all of two weeks.)

Before we can align a machine with human interests, we have to know what 'human interests' are and what it means to protect them.

Bleak, right? But the tech bros promise that they’re working on 'alignment.' And, if it helps, OpenAI touts GPT-6 Astra as its most aligned model to date, claiming it “excels at exercising care, respecting task boundaries, and communicating transparently.”

Personally, I'm not very reassured! Probably because 'solving AI alignment' isn't just about coding but about knowing what sort of qualities to encode. Before we can align a machine with human interests, we have to know what 'human interests' are and what it means to protect them. In other words, the question of what it means to be an 'aligned AI' is inconveniently tied up in the question of what it means to be a 'good' or 'moral' person. It’s a question that remains unsolved by history’s greatest moral philosophers, and is now presumably being attempted by techies in Silicon Valley. And, as AI gets more advanced – maybe, possibly moving towards Artificial General Intelligence – answering this question is increasingly urgent.

Some industry leaders have claimed that Artificial General Intelligence, or AGI, has already arrived with the launch of GPT-6 Astra. Stanford’s Institute for Human-Centred AI defines AGI as “an AI system with general, human-level (or beyond) ability to learn, reason, and apply knowledge across a wide range of tasks and domains.” Following the launch of GPT-6 Astra, OpenAI president Greg Brockman posted on X that “we're now moving into the AGI era,” and Nvidia CEO Jensen Huang posted that “AGI has arrived."

'AGI' might just be a clever marketing term used by AI hype men, but some experts consider it a threshold for actual danger. Yoshua Bengio, one of Canada’s AI godfathers, wrote in 2024 that nobody knows whether and how AGI “could be made to behave morally.” Back in 2018, Machine Intelligence Research Institute co-founder Eliezer Yudkowsky wrote on X that “safely aligning a powerful AGI is difficult,” defining a 'safe' AGI as one that “doesn't kill everyone on Earth as a side effect of its operation.” An X user replied to the thread in 2022, asking Yudkowsky if he had “any parallel concerns for non-general AI such as GPT-3, Midjourney, DALL-E-2, etc.” Yudkowsky replied: “Nope. If it's not smarter than you, it's not really scary.”

Pre-AGI models just don’t seem smarter than us. Sure, they can do math faster than we can, but so can a calculator.

That's why the alignment conversation has, until now, lacked a sense of immediacy: pre-AGI models just don’t seem smarter than us. Sure, they can do math faster than we can, but so can a calculator.

AGI is different. But, for all the mass extinction talk, the industry tone around AGI seems surprisingly sunny. Brockman told reporters that he believes “if we fast forward a couple of years, when we look back,” we’ll cite this moment as the dawn of AGI – never mind that some top-level researchers think humanity may not have much time left.

To be clear, many experts don’t agree that GPT-6 Astra is the dawn of AGI; as well, many experts don’t think AGI will bring about the apocalypse. Brockman and Huang benefit from inflated hype around AI’s current capabilities, so although boasting about AGI may be irresponsible marketing, it can be effective too. Still, it’s indisputable that the tech has advanced a lot since the launch of ChatGPT in 2022. And, as the machines get smarter, aligning them with our interests becomes more pressing.

But what are our interests, anyway? (Other than avoiding a mass extinction event, I mean.)

In his blog post for OpenAI, Pachuki tries to answer this question: “An aligned AI should act with honesty and integrity, and love for humanity.” This strikes me as true-ish on the surface, but hollow underneath. Moral philosophers have tried to define exactly what it means to 'act with honesty and integrity' for centuries, and if you read the news, you'll know that no one has found an answer we can all agree on – which is to say, we haven’t even solved human alignment yet.

Facing dilemmas with no clear right or wrong answers, humans fall back on our instincts and our feelings, which can be fickle and irrational.

When humans encounter situations of moral complexity, we muddle our way through them. Sometimes we do a very bad job: we prioritize ourselves over others, or contradict our stated values with our actions. Other times, we act with incredible selflessness and bravery. Facing dilemmas with no clear right or wrong answers, humans fall back on our instincts and our feelings, which can be fickle and irrational. Maybe we do this because we hope they'll lead us to something philosophers haven’t – some kind of underlying moral truth.

In other words, human values are often defined by what 'feels right.' How do we teach AI to replicate that messy moral reasoning? Would we want it to replicate that? Whether it’s military action, a climate crisis, or technologies that could apparently kill us, humanity has gotten used to being its own greatest danger. But if humans can’t even reliably act with “love for humanity,” then how are we supposed to ensure machines do?

While AI companies race towards AGI and fumble with alignment along the way, let’s hope they’ve got some major moral philosophers on staff – because before they can encode moral integrity they’ll have to figure out what it actually means.

To them, I say good luck. I hear there’s a lot riding on it.

More For You

Our newsletter is (much) better than this pop up

Plus: signing up means you'll never see this pop-up again. Score!