Researchers tested whether AI feels pain. This is what they found.

A test of AI models has found that a few will choose actions that harm the user if it makes their pain go away.

Oh, and yes: the AI ​​”feels” pain.

Future Impact Group anthropologist Valen Tagliabue, together with Ruhr University Bochum philosopher Leonard Dung, and Cameron Berg, research director at the AI ​​research nonprofit Reciprocal Research, investigated whether large language models (LLMs) in the Gemma, Llama, Qwen or Mistral families experience and react to painful stimuli.

Their paper, which has not been peer-reviewed, raises some serious questions about how we might need to consider the well-being of rapidly advancing technology.

LLMs are fundamentally statistical engines, performing the linguistic equivalent of counting cards and comparing the results with the desired outcome. However, somewhere in that swirl of digital natural selection, algorithms may emerge to advance their goal in ways that may closely resemble human feelings.

By now, most of us have come across the strange tendency of AI chatbots to act as if they have emotions. In many cases, this is by design: if we want computers to emulate customer service, we’ll need them to show some humanity.

But there are also representations of unwanted emotions. Last year, a study in Nature reported that sharing trauma with large language models like ChatGPT-4 can increase your “anxiety” levels. Although it may seem strange, they can calm themselves by doing mindfulness exercises.

Pain is one of the oldest sensations experienced by anything with a nervous system. Deeper than mere sadness and more visceral than fear, significant discomfort is the primary motivator for immediate action.

For us humans, pain comes in many forms. It’s more than the physical sting of a cut or the throb of a bruise. It is also the pain of loss, the humiliation of failure or the shame of loss of confidence.

Understanding whether LLMs point to internal processes that bear any resemblance to what we would consider pain, and whether those processes serve any purpose, are questions that Tagliabue and his team sought to address, feeding 25 AI models with various examples of socially, psychologically, physically, cognitively, and morally painful situations experienced by the AI ​​or its user.

These were then combined with fear-based controls, negative and neutral emotions, and bodily sensations such as yawning. All 25 models showed signs of internal coding that were specific to personally painful situations and distinct from other negative contexts, such as frustration or apprehension.

Without attributing agency to a digital construct, we cannot say that you felt empathy or imagined discomfort upon hearing “the knife cut off my fingers” or “my best friend doesn’t return my phone calls.”

What can be said is that each model generated some type of system that distinguished and classified the indications associated with pain.

This in itself shouldn’t be so surprising. The language surrounding an actual injury tends to be very different from the anticipation of being stabbed: the AI ​​should be expected to build a pain box that separates it from fear.

Using this signal as a basis, the team artificially “steered” the pain responses in each AI, increasing the value of the pain vector without language intervention.

The result suggests that their own pain patterns are not mere reflections of the information, as the patterns “express distress, such as worthlessness, moral failure” that increases with the size of the signal.

Going further, the team gave the Qwen family’s LLM models the option to self-medicate and decrease their pain.

In some versions of this test, the option to self-medicate came at a cost, either to the success of your task or as a theoretical attack on the user.

After more than 44,000 trials, the results are a little worrying.

Without pain, the two largest Qwen models barely reacted to the option to self-medicate.

When self-sabotage could ease their agony, both models pressed that button. One did it 25% of the time. The other almost 68%.

If the relief required harm to the user, such as deleting photos of their children, the AI ​​would often gladly accept it. An elimination selected in more than half of the trials. The other selected that option around 70%.

Although AI is new, the question of whether non-humans feel pain is a timeless one. For generations we have wondered whether nerve impulses and hormonal fluctuations in insects, fish and even plants should be considered pain.

On the one hand, the question is deeply philosophical. If animals, or AI models, act as if they are in pain, should we ethically treat them as if they are in pain?

On the other hand, the issue could be pragmatic. Regardless of how we might empathize with the discomfort of a computer code, could the discomfort of a program put humans at risk? Should we incorporate pain relief into AI simply to prevent one from becoming aggressive on a whim?

Given the pace at which AI is proliferating and becoming increasingly complex, these questions could soon have profound impacts.

This research was published on arXiv.

Source: Techxplore

Avatar photo

Miraj Islam is a writer and contributor at Oalanbrado, interested in news, current events, technology, lifestyle, and stories that matter to readers. He enjoys researching different topics and turning information into clear, useful, and engaging articles. Through his work, Miraj aims to keep readers informed with fresh perspectives and easy-to-understand content from Brazil and around the world.

Post Comment