By the WahLiao desk · Last verified 1 October 2026
Super Intelligence, the kind of AI behind ChatGPT, Claude and Gemini, comes with its own vocabulary, and about 40 words cover almost everything you will meet in the news, in a SkillsFuture course or at a phone shop. The core idea is short: a model is trained on huge amounts of data, a large language model predicts text one token at a time, and a hallucination is when it states something false with confidence. Everything else on this page builds on those four words. Where a definition comes from an official source, we say so: mainly the US National Institute of Standards and Technology (NIST) and Singapore’s IMDA and AI Verify Foundation.
The single most useful rule: when a salesperson, headline or app uses a term you do not know, look it up before you pay for it or believe it. Most jargon describes something ordinary.
AI glossary: quick facts
| Terms covered | 41, from agent to voice clone |
| Start with | Model, large language model, token, prompt, hallucination |
| Official glossaries | NIST CSRC glossary and NIST AI 100-3; IMDA and AI Verify Foundation frameworks |
| Singapore terms | SEA-LION (AI Singapore), AI Verify, Project Moonshot |
| Rule of thumb | 100 tokens is about 75 English words (OpenAI) |
| House style | “Super Intelligence” is WahLiao’s name for all AI; “superintelligence” (ASI) is a hypothetical future system |
The WahLiao Verdict
| Most useful word | Hallucination. Knowing it exists changes how you read every answer. |
| Most oversold | AGI. No agreed definition, no agreed date. |
| Most practical | Context window: it explains why long chats go off track. |
| Most local | SEA-LION, Singapore’s own open models for regional languages. |
| Safe to ignore | TOPS counts on a laptop sticker, unless you need on-device features. |
A to B
Agent: a Super Intelligence system that does not just answer but takes actions for you, such as booking, filling forms or running code, in several steps. IMDA’s Model AI Governance Framework for Agentic AI, launched in January 2026, describes agents as able to “take actions, adapt to new information, and interact with other agents and systems to complete tasks on behalf of humans”. See AI agents: what they do and the risks.
AGI (artificial general intelligence): a hypothetical system that could match people across most intellectual tasks, not just one. There is no agreed definition or test, and forecasts of when it might arrive vary widely. See AGI and superintelligence explained.
AI Verify: Singapore’s testing framework and open-source toolkit for checking whether an AI system behaves responsibly, run by the AI Verify Foundation, a not-for-profit subsidiary of IMDA.
Algorithm: a set of step-by-step instructions a computer follows. Your bank’s fraud check and a GPS route planner use algorithms; machine learning is a way of producing algorithms from data rather than writing every rule by hand.
Alignment: the work of making a model do what its makers and users intend, and not do harmful things. It covers both everyday behaviour (refusing to help with scams) and longer-term research. See Safety and risk, plainly.
API (application programming interface): a way for one program to talk to another. When a Singapore start-up builds a chatbot on top of GPT or Claude, it usually sends your question to the model maker through an API and pays per token.
ASI (artificial superintelligence): a hypothetical system far beyond human ability in almost every field. It does not exist. On WahLiao, “Super Intelligence” with capitals is simply our house name for today’s AI, not a claim that ASI has arrived; see why WahLiao says Super Intelligence.
Benchmark: a standard set of test questions used to score and compare models. Useful, but a high score on a benchmark does not guarantee good answers on your own task, and models can be tuned to the test.
Bias: systematic unfairness in a model’s outputs, often inherited from its training data, for example a CV-screening tool that favours one group.
C to D
Chatbot: a program you talk to in plain language. Today’s chatbots, such as ChatGPT, Claude and Gemini, run on large language models; older ones followed fixed scripts. Singapore’s OneService chatbot on WhatsApp and Telegram is a government example.
Context window: how much text a model can take into account at once, including your prompt, pasted documents and the conversation so far, measured in tokens. When a long chat exceeds it, earlier parts drop out or get less attention, which is why very long chats drift.
Deepfake: synthetic video, image or audio that makes a real person appear to say or do something they did not. Scammers use them to pose as public figures promoting fake investments, or as someone you know. See deepfake and voice-clone scams and deepfakes and the law.
Diffusion model: the kind of model behind most image and video generators. It learns to turn random noise into a picture step by step, guided by your text description.
E to G
Embedding: a list of numbers that represents the meaning of a word, sentence or image, so that similar meanings sit close together. Embeddings let a search find “HDB flat” when you typed “public housing”.
Fine-tuning: further training of an existing model on a smaller, specialised set of data. NIST defines it as a training step that starts from a pre-trained model and “adds task- or domain-specific information”, for example a model tuned on a company’s customer-service replies.
Foundation model: a large model trained on broad data that can be adapted to many tasks. GPT, Claude, Gemini, Llama and SEA-LION are all foundation models. See who makes the models.
Generative AI: Super Intelligence that creates new content, text, images, audio, video or code, rather than only sorting or predicting from data. NIST’s Generative AI Profile describes it as models that “emulate the structure and characteristics of input data in order to generate derived synthetic content”.
GPU (graphics processing unit): a chip first built for video games that is very good at the parallel arithmetic models need. Large data centres of GPUs train and run today’s models, which is why chips and electricity feature in AI news. See energy, water and data centres.
Guardrails: rules and filters placed around a model to block harmful, off-topic or leaked content, such as refusing to write malware or reveal personal data. They reduce risk but can be bypassed.
H to L
Hallucination: when a model states something false as if it were true, such as a made-up citation or a wrong CPF rule. NIST calls this “confabulation”: “the production of confidently stated but erroneous or false content”. Singapore’s CSA and IMDA warn that such content is presented “in a convincing and authoritative manner”. See How Super Intelligence works.
Inference: the stage where a trained model is actually used to produce an answer. Training happens once and costs a fortune; inference happens every time you ask a question.
Large language model (LLM): a model trained on vast amounts of text to predict the next token, which lets it write, summarise, translate and answer questions. It is the engine inside ChatGPT, Claude, Gemini and SEA-LION.
Machine learning: the branch of Super Intelligence in which computers learn patterns from examples instead of following hand-written rules. Spam filters, recommendation feeds and large language models are all machine learning.
M to O
Model: the trained system itself, a very large set of numbers that turns inputs into outputs. A product such as ChatGPT may offer several models, faster or more capable, under one app.
Multimodal: able to handle more than one kind of input or output, for example reading a photo of a hawker menu and answering in text, or speaking aloud.
Neural network: a model built from layers of simple connected units, loosely inspired by the brain, whose connection strengths are adjusted during training. Almost all modern Super Intelligence uses neural networks.
NPU (neural processing unit): a chip on a phone or laptop designed to run AI features locally and efficiently. Microsoft requires an NPU capable of 40+ TOPS (trillion operations per second) for a Copilot+ PC. See do you need an AI phone or laptop.
Open-weight: a model whose trained numbers (weights) are published so anyone can download and run it, like Meta’s Llama or AI Singapore’s SEA-LION. This is not always the same as fully open-source, which would also include training data and code.
P to R
Parameter: one of the adjustable numbers inside a model that training sets. Large models have billions of them; more parameters usually means more capability and more cost, but not always better answers.
Prompt: what you type or say to a model. Clear context, an example and the format you want make a big difference; see using it well.
Prompt injection: an attack that hides instructions inside content a model reads, such as a web page or email, to hijack it. NIST describes it as exploiting the mixing of untrusted input with a trusted prompt. It matters most for agents that browse and act for you.
RAG (retrieval-augmented generation): a set-up in which the model first looks up relevant documents, then answers using them. NIST notes RAG lets a model draw on new information “without the need for retraining”; a bank chatbot that answers from its own FAQ pages is a typical example.
Reasoning model: a model trained to work through a problem in steps before giving its final answer, trading speed for accuracy on maths, logic and coding. It still makes mistakes.
Red-teaming: deliberately attacking a model to find ways it can fail or be misused, before the public does. Singapore’s Project Moonshot, from the AI Verify Foundation, is an open-source toolkit for benchmark testing and red-teaming large language model apps.
Reinforcement learning from human feedback (RLHF): a training step in which people rate or rank a model’s answers and the model is adjusted towards the preferred ones. It is a large part of why chatbots sound polite and helpful.
S to Z
SEA-LION: a family of open models from AI Singapore built for Southeast Asian languages and contexts, available free for research and commercial use. Its newest versions are developed with NVIDIA. See Singapore’s own AI story.
Synthetic data: data generated by a computer, often by another model, rather than collected from the real world. It is used to train or test models when real data is scarce or private, but can carry forward the errors of the model that made it.
Token: the small chunk of text a model reads and writes, often part of a word. OpenAI’s rule of thumb for English is one token for about four characters, or 100 tokens for about 75 words. Prices for developers and limits on context windows are counted in tokens.
Training: the process of feeding a model huge amounts of data and adjusting its parameters until its predictions improve. Pre-training builds general ability; fine-tuning and RLHF shape it afterwards. See the history of AI.
Transformer: the neural network design, introduced by Google researchers in 2017, that underlies almost every large language model today. NIST describes GPT models as “based on the transformer architecture”. It lets a model weigh how every word in a passage relates to every other.
Voice clone: a synthetic copy of a real person’s voice made from a short recording. Scammers use them to pose as family members on the phone; agree a family code word and call back on a number you know. See how to check.
AI glossary: FAQ
What is the difference between AI, machine learning and a large language model?
AI is the broad field, machine learning is the main way it is built today (learning from data), and a large language model is one kind of machine learning model, specialised in text.
What does AI hallucination mean?
It means a model stated something false with confidence, such as an invented source or wrong figure. NIST’s formal term is confabulation. Always check important answers against the original source.
What is a token in ChatGPT?
A token is a chunk of text, often part of a word. In English, 100 tokens is roughly 75 words, according to OpenAI. Context windows and developer prices are measured in tokens.
Is AGI the same as superintelligence?
No. AGI means roughly human-level ability across most tasks; artificial superintelligence (ASI) means far beyond it. Neither exists today, and WahLiao’s “Super Intelligence” label is a house name for ordinary AI, not a claim about either.
What is SEA-LION?
SEA-LION is a family of open language models from AI Singapore designed for Southeast Asian languages and cultures, free to download for research and commercial use.
Read next
This page belongs to Super Intelligence. Next, read Who makes Super Intelligence.
Sources checked 1 October 2026: NIST CSRC glossary, fine-tuning; NIST CSRC glossary, retrieval-augmented generation; NIST CSRC glossary, prompt injection; NIST CSRC glossary, generative pre-trained transformer; NIST AI 600-1, Generative AI Profile (July 2024); NIST AI 100-3, The Language of Trustworthy AI glossary; CSA and IMDA, Safe & Secure Use of Generative AI for Individuals; Allen & Gledhill, IMDA Model AI Governance Framework for Agentic AI; AI Verify Foundation, Project Moonshot; EDB, AI Verify Foundation; AI Singapore, SEA-LION; MDDI, OneService chatbot; OpenAI Help Center, what are tokens; Microsoft, Copilot+ PC NPU requirement. General information. Last updated 1 October 2026.
