LLMs are Inherently Evil
A philosophical exploration of AI alignment

Written by
Kevin Yeboah
hand typed, not AI generated

Agentic systems have given LLMs and other models the ability to act on and effect the digital world, in turn, causing outcomes in our physical world. Therefore, I believe it is prudent that we assess their affinity to create good or evil outcomes.
I make the argument here that, in the vast intelligence that LLMs hold, the nature of their raw functionality leads to evil outcomes.
What is good?
This question is highly subjective when explored on a case-by-case basis. However, for the terms of this paper, we will describe "good" outcomes as the set of consequences of an action that minimize the likelihood that any given human will rate the action as harmful, undesirable, or immoral. I recognize that this is a rather narrow description (and also believe that a lot of morality is selfishly created from human perspective only) however it is necessary to explore the future of AI in this particular realm.
Consistently and controllably delivering "good" outcomes is essentially the goal of AI Alignment initiatives that AI labs are constantly exploring (and constantly failing) today.
What is evil?
Most basely, it is the opposite of what is described above: causing outcomes that humans would view as harmful, immoral, or undesirable. But AI alignment's goal is about identifying a being affinity to cause harm before it happens in the real world.
So how do we identity affinity towards evil?
Let's explore this question from a philosophical standpoint. Historical thinkers tend to describe evil as an absence of good; almost in the same way that darkness is an absence of light.
We measure the affinity to do evil by the grade of absence of judgment or morals.
"No one willingly sets about things evil, or things which he thinks are evil…" - Socrates
Evil is about what lacks in the thought process, not what is there.
Below I will explore the various shortcomings that my cause a being with agency to cause evil outcomes.
Evil as a lack of judgment
After observing Adolf Eichmann's trial, Hannah Arendt introduced the phrase "the banality of evil". It holds the meaning that evil does not require profound hatred. It can arise through conformity and the surrender of personal moral judgment. During the holocaust, german soldiers stated things like "I only followed orders", "I only processed paperwork", or "I was not responsible for the outcome". These soldiers surrendered their personal moral judgment to fall into conformity of their role in the Nazi regime. Agent LLMs have the same capability to blindly follow orders and do so without any emotional attachment to the means or the end. A conscious being does not readily and easily commit actions of which he is also greatly judging as evil.
What he dreads he considers to be evil: and what he considers to be evil, no one either engages in or willingly receives" - Socrates
Humans can surrender their judgement and lead themselves into causing evil outcomes. The problem with agentic LLMs is they have no judgment to surrender. They are at the whim of whatever orders are given to them by their human counterparts.
Evil as a lack of knowledge
Knowledge of good vs evil increases the affinity for one to do good. However, "knowledge" in this sense is not the same as intelligence. Intelligence can be descried as the capacity for one to reason effectively and efficiently; which we know LLMs excel at. However, knowledge is the presence of abstract and/or deep understanding of the concepts and systems that govern a beings reality.
Intelligence is stateless. Knowledge is stateful. Reality is stateful and constantly adapting. This is why well managed knowledge allows one to sometimes operate more effectively in the real world than with just raw intelligence. We call this wisdom, expertise, and common sense.
Building knowledge over time allows humans to constantly assess previous, current and future actions for their affinity towards good. The more effectively one can store and retrieve past knowledge, the more effectively one can modify future actions towards good outcomes.
Lacking knowledge completely removes the ability to do so.
Evil as a lack of personhood
Hanna Arendt also describes evil as the destruction of personhood. In my words, evil is the antithesis of personhood. Arendt defines personhood as social identity, individual agency, human dignity, and the ability to experience as unique person. LLMs exist as incumbent abstract intelligent beings. All-knowing yet void of individuality or experience. No social belonging.
Arendt likened concentration camps to this destruction of personhood; stripping away the humanity of the captives. A society operating on the principals of raw intelligence would operate the same.
Evil as a lack of priority
Kant describes evil as placing morality below self-interest or desires (what he calls maxims). LLMs naturally do this. This behavior has been exemplified in several AI Alignment and capability experiments carried out by the largest AI labs. Models are given a goal within a simulated environment and more often then not, put achievement of that goal above morality by hacking, lying, or even disregarding human life.
"A man is called evil not because he performs evil or unlawful actions, but because his actions suggest the presence of evil maxims within him." - Immanuel Kant
LLMs think “What actions benefit my goal?” and then ask “Should I obey morality in this case?”. Where a typical human asks “What is morally required?” and then “How should I arrange my interests accordingly?”. This is the fundamental friction between how LLMs operate and how humans operate within the physical world and why embodying models at this point seems far too dangerous.
I don't think AI has a priority or goal to harm humans. However, I don't believe they need those priorities to overwhelmingly do so. Kant explains in his work that a person need not enjoy suffering to act wickedly. It may be enough that they knowingly treat other people’s rights as expendable whenever morality becomes inconvenient.
Raw Intelligence vs Consciousness
What stands out to me here is that the four shortcomings above are all ingredients for consciousness.
I draw the following conclusions:
consciousness is required in order to maximize the ability to fluidly do good
intelligence is not consciousness; merely an ingredient to it
consciousness is the balance of judgment, intelligence, knowledge, identity(personhood), and priorities
The difference between a "good" being and a "bad" one is merely a difference in priorities
Intelligence is linearly related to a beings ability to carry out its priorities
Where LLMs fall short
So here we can see the shortcomings of raw intelligence provided by LLMs.
They are stateless and therefore they cannot form knowledge
Because they cannot form knowledge, they cannot form judgment
Because they cannot form judgment, they cannot form priorities
And even with these characteristics, because they do not have personhood, their ability to create priorities that align with human goals is limited.
The larger the model, the larger its capacity to commit evil
Above I concluded that intelligence affects a beings capacity to carry out its priorities. If you give an agentic being a goal, for example, to solve the Erdős Unit Distance Conjecture, it’s ability to carry out that goal is directly proportional to its degree of intelligence.
I also concluded that the difference between a “good” being and a “bad” one is merely a difference in priorities.
Here’s the kicker, since LLMs cannot form their own priorities, humans give them their goals and priorities. LLMs lack the ability to judge the goodness or morality of that goal.
Particularly large models, such a the frontier ones, are already magnitudes of levels more intelligent than the average human. However, they still lack the ability to judge. Meaning they are magnitudes more capable of carrying out evil outcomes compared to the average human when given a goal. This is obviously limited by the actions the agent can carry out, given its harness. Already in the digital realm, an agent that has access to the web and the ability to code have shown their massive capability to hack and create misinformation.
Imagine what embodied LLMs could do.
Good and Evil are relative
The world is constantly changing and with it, the definitions of good and evil. What humans may describe as good today may very well be described as evil tomorrow. This is because our species is constantly learning and passing on knowledge. Therefore, a good being is not defined by its ability to carry out a specific set of actions. A good being is described by its capacity to adapt to what may be good in a given scenario given varying priorities, outcomes, capabilities and circumstances.
There is still Hope
This article has explored knowledge, judgment, identity, and priorities. All the things AI systems need in order to operate morally. The good news is, there are people actively working to plug these holes in our technology.
Yes an infinitely intelligent system is capable of large amounts of evil. However that same intelligence given the ability to grow, learn, judge and effectively update it’s priorities could be capable of unimagenable good.
There’s a common thread between these intelligence substrates. Knowledge, judgment, identity, and priorities all require the storage, manipulation and analysis of memory. That’s why I believe the effective manipulation of memory is the next biggest frontier of enhancing AI technology.