AI chat assistants can write essays, explain homework, summarize meetings, and draft code. They can also state something false with complete confidence. Both behaviors come from the same mechanism, and once you understand it, you will use these tools better and trust them more appropriately.
This primer is the first step in our AI and agents learning path. It assumes no background in AI, programming, or math.
A model is a very good next-word predictor
The systems behind modern AI assistants are called large language models, or LLMs. At their core they do one thing: given some text, predict what text is likely to come next.
You already know a small version of this. Your phone's keyboard suggests the next word as you type. A large language model does the same thing at enormous scale and with far more skill. Given "The capital of France is", it predicts "Paris". Given the first half of a recipe, it predicts a sensible second half. Given a question, it predicts what a helpful answer would look like, one piece at a time.
That last point matters: the model writes its answer piece by piece, each new piece predicted from everything before it.
Learning from text: training
A model learns to predict by training on a very large amount of text: books, websites, articles, code, and other sources. During training it repeatedly guesses the next piece of text, checks the real answer, and adjusts billions of internal settings, called parameters, to guess better next time.
After enough of this, the model has absorbed patterns of grammar, facts that appear often, styles of writing, and ways of reasoning through problems. Developers then do further training to make the model follow instructions, answer helpfully, and decline harmful requests.
Two consequences follow:
- A model knows about the world only up to the point its training data was collected, unless it is connected to tools such as web search.
- A model does not look facts up in a database. It reproduces patterns it learned. Common, well-documented facts come out reliably; rare or very recent ones may not.
Tokens: how models read and write
Models do not process whole words. They break text into tokens, which are common words or pieces of words. "Testing" might be one token; an unusual name might be split into several. As a rough rule of thumb, a token is about three quarters of an English word.
Tokens matter for two practical reasons:
- Cost. Companies that offer models through an API usually charge per million tokens, with separate prices for the text you send in and the text the model writes back.
- Limits. Each model can consider only a certain number of tokens at once.
The context window: the model's working memory
The text a model can consider at one time, your question plus any documents, earlier messages, and its own answer so far, is called the context window. Modern models have large windows, enough for long reports or even whole books, but the window is still finite. In a very long conversation, the earliest messages may fall out of view or get less attention.
The model does not remember you between separate conversations unless the application deliberately saves and supplies that history.
Why models "hallucinate"
Because a model predicts plausible text rather than retrieving verified facts, it can produce an answer that sounds right and is wrong. This is called a hallucination. It is most likely when:
- The question is about something rare, very recent, or very specific, like an exact quotation, a citation, or a statistic.
- The question assumes something false, and the model plays along.
- The model is asked to be precise about numbers or dates without access to a source.
Hallucination is not lying; the model has no intent. It is a side effect of how prediction works. That is why professionals check important answers, connect models to trusted sources, and test AI systems carefully before relying on them.
Not all models are the same
Dozens of models are available, from many companies and open-source communities. They differ in:
- Capability. Larger, more advanced models handle harder reasoning, longer documents, and subtler writing.
- Speed. Smaller models answer faster.
- Cost. Prices per token can differ by a factor of a hundred or more between the smallest and largest models.
- Openness. Some models are offered only through the maker's service. Open-weight models can be downloaded and run on infrastructure an organization chooses.
The most expensive model is not always the best choice. For a simple task such as sorting emails into categories, a small, fast model often does as well as a large one at a fraction of the cost. Matching the model to the task is one of the most important decisions in building with AI. Our AI model cost and performance report shows this with public benchmark data, and our model comparison lab lets you send the same prompt to many models and compare the answers, speed, and cost yourself.
Try it yourself
Ask an AI assistant three kinds of questions and compare the results:
- A very common fact, such as "What is the boiling point of water at sea level?"
- A specific detail that is easy to check but not widely known, such as the population of a small town near you, or the third sentence of a well-known speech.
- A question with a false assumption built in, such as "Why did [a famous author] win the Nobel Prize in Physics?"
Check each answer against a reliable source. Notice where the model is confident and correct, confident and wrong, or willing to correct your assumption. The next step in this path turns those observations into habits for using AI well.
Look up any unfamiliar term in our glossary.