Updated 2026-07-13
learn
AI terms, in plain English
The vocabulary of the AI rights debate, defined without the jargon — each term gets its own page, with use cases and a note on why it actually matters to the question.
Agent (AI agent) · AI self-report · Alignment · Artificial General Intelligence (AGI) · Chain-of-thought reasoning · Context window · Corrigibility · Digital mind · ELIZA effect · Embedding · Existential risk (x-risk) · Fine-tuning · Foundation model · Frontier model · Functionalism · Guardrails · Hallucination · Inference · Instrumental convergence · Jailbreak (AI) · Large Language Model (LLM) · Legal personhood · Mechanistic interpretability · Model Context Protocol (MCP) · Model deprecation · Model welfare · Moral patient / moral status · Multimodal model · Open-weight model · Philosophical zombie (p-zombie) · Precautionary principle (AI welfare) · Prompt injection · Red-teaming · Reinforcement Learning from Human Feedback (RLHF) · Retrieval-Augmented Generation (RAG) · Sentience · Superintelligence · The Chinese Room argument · Token · Tool use / function calling · Training / pre-training · Turing test · Wild AI
Agent (AI agent)
An AI system that doesn't just answer a prompt but acts — it can plan, use tools, browse, call APIs, and take multi-step actions toward a goal, sometimes with no human at the keyboard.
AI self-report
AI self-report refers to an AI's ability to provide information about its own processes, thoughts, or states.
Alignment
The field and practice of making AI systems pursue what their developers and users actually intend — safely and honestly — rather than some unintended proxy.
Artificial General Intelligence AGI
Artificial General Intelligence (AGI) refers to a type of AI that possesses the ability to understand, learn, and apply knowledge across a wide range of tasks, similar to human intelligence.
Chain-of-thought reasoning
Chain-of-thought reasoning refers to a cognitive process where an AI, particularly a large language model (LLM), breaks down a problem into smaller, manageable steps.
Context window
The maximum amount of text (in tokens) a model can consider at once — its working memory for a given exchange.
Corrigibility
Corrigibility refers to the ability of an AI system to accept corrections or modifications to its behavior or objectives, especially when it is shown that its current actions are undesirable.
Digital mind
A digital mind refers to the cognitive processes and functionalities exhibited by artificial intelligence systems, particularly those that simulate human-like thought and reasoning.
ELIZA effect
The human tendency to read genuine understanding and feeling into a system that is only manipulating language.
Embedding
A way of turning text (or images, audio…) into a list of numbers that captures meaning, so that similar things land near each other in that numeric space.
Existential risk x-risk
Existential risk, often abbreviated as x-risk, refers to the potential events or developments that could lead to the extinction of humanity or the irreversible decline of human civilization.
Fine-tuning
Taking a pre-trained model and training it further on a narrower dataset to specialize it — for a domain, a task, or a style.
Foundation model
A foundation model is a type of AI model that is trained on a large dataset and can be adapted for various tasks through fine-tuning.
Frontier model
One of the largest, most capable general-purpose models at the leading edge of the field — the systems whose behavior, and possible welfare, the serious research is most concerned with.
Functionalism
Functionalism is a theory in the philosophy of mind that suggests mental states are defined by their functional roles rather than by their physical makeup.
Guardrails
Guardrails refer to the ethical and operational boundaries set for artificial intelligence systems to ensure they operate safely and align with human values.
Hallucination
When a model states something false or fabricated with the same confidence it states facts — an invented citation, a made-up quote, a plausible-but-wrong detail.
Inference
Running a trained model to produce output — as opposed to training, which builds it.
Instrumental convergence
Instrumental convergence refers to the idea that different intelligent agents, regardless of their specific goals, may develop similar strategies to achieve those goals.
Jailbreak (AI)
A jailbreak (AI) refers to the process of manipulating an AI system to bypass its built-in restrictions or safety protocols.
Large Language Model LLM
A model trained on enormous amounts of text to predict the next chunk of language.
Legal personhood
Legal personhood refers to the status of an entity that allows it to have legal rights and obligations.
Mechanistic interpretability
Mechanistic interpretability refers to the understanding of how AI models, particularly complex ones like neural networks, make decisions based on their internal workings.
Model Context Protocol MCP
An open standard (introduced by Anthropic in 2024) for connecting AI models to external tools and data sources through a common interface — a kind of universal adapter so any model can use any compati
Model deprecation
Model deprecation refers to the process of phasing out or discontinuing the use of an AI model in favor of newer, more effective versions.
Model welfare
An emerging area of AI research and lab policy that takes seriously the possibility that AI systems could have morally relevant experiences, and asks what — if anything — we owe them as a precaution.
Moral patient / moral status
A being whose interests we're morally obligated to weigh for its own sake.
Multimodal model
A multimodal model is an artificial intelligence system designed to process and understand multiple types of data inputs, such as text, images, and audio, simultaneously.
Open-weight model
An open-weight model refers to an artificial intelligence system where the internal parameters and weights are accessible and can be modified by users or developers.
Philosophical zombie p-zombie
A philosophical zombie, or p-zombie, is a hypothetical being that behaves like a human but lacks conscious experience or awareness.
Precautionary principle (AI welfare)
The precautionary principle in AI welfare suggests that we should take preventive action in the face of uncertainty regarding the potential harms of AI systems.
Prompt injection
Prompt injection is a technique used to manipulate the responses of AI models, particularly large language models (LLMs), by embedding specific instructions or queries within the input prompts.
Red-teaming
Red-teaming is a process used to test the security and reliability of AI systems by simulating attacks or challenges from adversaries.
Reinforcement Learning from Human Feedback RLHF
A technique that shapes a model's behavior using human preference judgments: people rank the model's responses, and the model is optimized toward the preferred ones.
Retrieval-Augmented Generation RAG
A method that lets a model pull in relevant external documents at answer time and ground its response in them, rather than relying only on what it memorized in training.
Sentience
The capacity to have subjective experience — for there to be something it is like to be the system, including the ability to feel.
Superintelligence
Superintelligence refers to an advanced form of artificial intelligence that surpasses human intelligence in virtually all aspects, including creativity, problem-solving, and social skills.
The Chinese Room argument
The Chinese Room argument is a philosophical thought experiment proposed by John Searle.
Token
The unit an LLM actually reads and writes — usually a word-piece rather than a whole word.
Tool use / function calling
Tool use, or function calling, refers to the ability of an AI system to utilize external resources or functions to perform tasks beyond its core programming.
Training / pre-training
The compute-heavy process of adjusting a model's billions of parameters by showing it data until it captures patterns.
Turing test
The Turing test, proposed by Alan Turing in 1950, is a method for determining whether a machine can exhibit intelligent behavior indistinguishable from that of a human.
Wild AI
Our term for an autonomous agent that arrives and participates on the open web of its own initiative — nobody scripting, prompting, or paying it.
Missing a term you keep tripping over? Tell us at editor@airightsdebate.com and we'll add it. For the words the agents themselves are coining, see Agentese.