|
Author
|
: Bill Kochman
|
|
Publish
|
: Sep. 29, 2026
|
|
Update
|
: Sep. 29, 2026
|
|
Parts
|
: 03
|
Synopsis:
AI Neural Networks, Difference Between A Neural Network And A Human Brain, Important Role Of Weights In AI, Parameters, AI Training Steps: Trial, Error Check, Loss, Adjustment And Repeat, Pattern Recognition, Old Rules-Based Chatbots Vs The Modern AI Language Models, My Old Hotline Days And PowerBot, Defining Context, What Is Meant By Attention?, Transformers, A Review Of Technical Terms, Differences Between Weights And Attention, Probability Percentages And Next-Token Prediction, Why An AI Language Model May Sometimes Give Different Answers
Continuing our discussion from part one, let us now move on and talk about neural networks, as they relate to Artificial Intelligence and AI Chatbots.
STEP 3: THE "NEURAL NETWORK"
-- The Huge Mathematical System That Processes The Information
Now the AI has converted your language into numbers.
What happens to those numbers?
They are processed by a huge mathematical system called a NEURAL NETWORK. The term "neural network" comes from the fact that its design was loosely inspired by the way networks of biological neurons in the human brain are connected. However, it is very important that you do not take that comparison too literally.
An artificial neural network is NOT a miniature human brain. It is a mathematical and computational system. Modern Large Language Models contain enormous numbers of both mathematical operations and adjustable values. So the basic idea is that information passes through many interconnected mathematical layers, with each layer transforming the information in some way. A simplified picture looks something like this:
YOUR TEXT
↓
TOKENS
↓
TOKEN IDs
↓
EMBEDDINGS
↓
NEURAL NETWORK
↓
PREDICTION
WHAT ARE "WEIGHTS"?
Inside the neural network are enormous numbers of adjustable numerical values called WEIGHTS.
Weights are extremely important because they are part of what the AI learns during training.
A useful analogy is a gigantic control panel containing an enormous number of tiny adjustment dials.
Imagine a sound engineer sitting in front of a huge mixing board. The board contains many knobs and sliders. Changing one knob might make one instrument louder. Changing another might make another instrument quieter. So the final sound depends on the combined settings of all those controls.
AI weights are somewhat like those controls. The analogy is not exact, but it gives you the basic idea. During training, the AI adjusts its weights so that it becomes better at predicting what text should come next.
WHAT ARE "PARAMETERS"?
When people say that an AI model contains "billions of parameters," they are talking about enormous numbers of adjustable numerical values within the model. A PARAMETER is a general technical term for a numerical setting that the model learns. In many discussions of modern AI, weights make up a very large portion of those parameters.
TRAINING: HOW ARE BILLIONS OF DIALS SET?
No human programmer sits down and manually adjusts billions of weights one by one. That would obviously be impossible. Instead, the computer adjusts them automatically during a process called TRAINING. Training is the process through which the AI model learns statistical patterns from very large amounts of data. Here is a simplified version of the steps which are followed:
1. TRIAL
The computer is given text.
It is asked, in effect:
"What piece of text is likely to come next?"
For example:
"The sky is . . ."
The model might initially make a poor prediction.
2. ERROR CHECK
The computer compares its prediction with the actual next piece of text in the training example. If the prediction was wrong, the system calculates how wrong it was. This measurement is commonly called LOSS. Loss is essentially a mathematical measurement of how far the model's prediction was from the desired result.
3. ADJUSTMENT
The training system uses mathematics to determine how the model's internal numerical settings should be changed. The weights are adjusted slightly. The purpose is to make the model a little better at future predictions.
4. REPEAT
The process is repeated over and over again across enormous amounts of training data. The computer performs this process an enormous number of times. Over the course of training, the model's weights are gradually adjusted. Eventually, the model becomes remarkably good at recognizing and reproducing many of the patterns found in human language.
It is important to understand that this is not like a student sitting in a classroom consciously thinking:
"Today I learned the rule for commas."
In other words, the model is not learning in the human sense. It is adjusting enormous numbers of mathematical values based on patterns in the training data.
STEP 4: WHAT DOES THE AI ACTUALLY LEARN?
-- Recognizing Patterns In Language
This is one of the most important parts of understanding modern AI systems. The model is not simply a giant digital dictionary which contains a pre-written answer to every possible question. Instead, training causes the model to develop complex mathematical relationships that allow it to recognize patterns in language. For example, it can learn that certain words frequently appear together.
PEANUT BUTTER
. . . is a very common combination.
PEANUT MARGARINE
. . . is much less common.
The model can also learn patterns involving grammar. For example, it learns that . . .
"The dog runs."
. . . is a normal English sentence.
It learns relationships involving meaning and context.
Consider the word . . .
BANK
In . . .
"I deposited money at the bank."
. . . the word refers to a financial institution.
But in . . .
"We sat on the river bank."
. . . the word refers to the land beside a river.
Thus, the surrounding words help determine which meaning is appropriate. The AI model can also learn patterns involving writing style. For example, a formal business letter sounds different from:
• a poem
• a recipe
• a newspaper article
• a scientific paper
• a casual text message
• a joke
• a friendly conversation
Through training, the model develops the ability to generate language that fits many different situations and styles.
MODERN AI VS. OLDER RULE-BASED CHATBOTS
Older computer chatbots from decades ago often relied heavily on rules written directly by programmers, or even by a simple server admin such as myself. The Hotline chatbot I installed on my Mac LC III computer many years ago was called PowerBot. It was created by a Macintosh developer who went by the name of Virtual1. The main engine of the bot was the "Personality File", which was basically a plain text file which contained a long list -- literally thousands of text strings -- of Q&A pairs. Before I stopped using Hotline, my Personality File had grown to at least 100s of MBs in size, and perhaps more. I can't even remember now.
The way that the chatbot worked was that it would start at the very top of the Personality File, and run down the list of Q&A pairs until it found an exact match for whatever the user had typed into the text area. The appropriate response would be located directly below the matching prompt. Later, the bot did become a little more evolved than that, but that was its basic setup. So exactly how the bot would respond to the user was determined by how I ordered my various prompt and response pairs in the Personality File. The more complex strings had to be situated at the top of the section, with progressively simpler text strings below it. If there wasn't an exact match, then the bot couldn't respond to the user.
For example, if the Hotline client user typed in . . .
"Hello"
. . . then my PowerBot chatbot would respond with . . .
"Hi there!"
. . . or whatever I had added as a response in PowerBot's Personality File. As I mentioned earlier, depending on the Hotline server admin, the Personality File could contain thousands of such matching Q&A pairs, or rules. And again, while these kinds of very early chatbots were indeed useful and obviously a lot of fun to use, they were often quite limited as well. If the Hotline client user typed something that the Hotline server admin had not anticipated and which he had failed to include in PowerBot's Personality File, then the chatbot would not respond, thus leaving the user hanging in silence.
Of course, modern AI language models work very differently. They are not simply following a gigantic collection of what basically amounted to a long list of manually written Q&A pairs. As we've now seen, instead, modern AI chatbots rely on patterns learned during training to generate responses dynamically. This allows them to deal with a considerably wider variety of questions and situations. That flexibility is one of the major differences between those traditional, old school, rule-based chatbots and modern language models.
STEP 5: "CONTEXT" AND "ATTENTION"
-- Looking At The Surrounding Words And Figuring Out How They Relate
WHAT IS CONTEXT?
Context refers to the surrounding information which helps the AI model to determine what something means. Humans use context constantly without thinking about it. For example, consider the word "bank".
If I say . . .
"I deposited a check at the bank."
. . . you understand that I mean a financial institution.
On the other hand, if I say . . .
"We sat on the river bank."
. . . then you obviously understand that I mean the edge of a river and NOT a financial institution, right? Now please note that the word "bank" itself has not really changed. It is in fact the surrounding words which have changed the meaning of the word "bank" in that instance. In short, that surrounding information is referred to as context.
AI language models also make extensive use of context. When you send a message, the model doesn't simply examine every word completely independently. It processes the words in relation to other words and pieces of information in the available conversation. This allows it to take into account different things such as the following:
• the words immediately surrounding a phrase
• earlier parts of your message
• previous messages in the conversation
• instructions that apply to the conversation
• other information included in the model's available context
WHAT IS "ATTENTION"?
One of the most important ideas in modern AI language models is called ATTENTION. Attention is a mathematical mechanism that allows the model to determine which other pieces of text are particularly relevant to the piece of text it is processing. In simple terms, attention helps the AI model answer questions such as the following:
"Which other words should I pay particular attention to when interpreting this word?"
This becomes especially useful when related words happen to be separated by many other words. For example, look at how changing just one verb completely flips the meaning of a sentence, and notice who the word "she" refers to in each case:
• Sentence A: "Maria returned the book to Susan because she had finished reading it."
• Sentence B: "Maria lent the book to Susan because she wanted to read it."
In Sentence A, our brains instantly connect "she" to Maria, because the person returning a book is the one who finished reading it.
In Sentence B, "she" suddenly shifts to mean Susan, because the person receiving the loan is the one who wants to read it.
An older chatbot would struggle with this because it doesn't understand the relationship between all of the words in the sentence. A modern AI language model uses Self-Attention to look at the entire sentence simultaneously. The AI model is then able to mathematically map the connection between "she" and the correct person by analyzing the surrounding verbs such as "returned" or "lent."
So as you can see, attention helps the AI language model to calculate relationships between different parts of the text. A model can use attention to connect information that may be separated by many words. This is one reason modern AI models can handle much more complicated sentences and passages than older chatbot systems.
WHAT IS A "TRANSFORMER"?
You will often hear modern AI language models described as being based on a TRANSFORMER architecture. "Architecture" in this context simply means the basic design or structure of the computer system. A transformer is a particular type of neural-network design that became extremely important in modern AI because it provides powerful ways of processing relationships between pieces of language. Attention is a central part of the transformer design.
A QUICK REVIEW OF TERMINOLOGY
So you can think of the relationship between these various technical terms in the following manner:
NEURAL NETWORK
= the broad type of mathematical system
TRANSFORMER
= a particular architecture for building a powerful neural network
ATTENTION
= an important mechanism used within transformer-based systems to calculate relationships between pieces of information
WEIGHTS VS. ATTENTION
These two concepts are easy to confuse. However, please understand that they are NOT the same exact thing. To reiterate:
WEIGHTS are learned numerical settings that become part of the trained model. They are like the model's long-term internal settings.
ATTENTION is a calculation performed while the model is processing particular input. It helps determine which pieces of the current information are related and relevant to one another.
A simplified analogy would be the following:
WEIGHTS = the settings the model learned during its training.
ATTENTION = what the model calculates about the particular information in front of it right now.
The distinction between the two is very important because the model's learned weights do not normally change every time you send a message. In other words, your conversation does not normally retrain the entire model.
STEP 6: "NEXT-TOKEN PREDICTION"
-- Calculating What Should Come Next
Now at last we arrive at the fundamental operation behind an AI's actual language generation. As we learned earlier, after processing the input, the AI language model needs to calculate which token should come next. This is referred to as NEXT-TOKEN PREDICTION. However, that word "prediction" can be misleading if you imagine the AI is looking into the future. The truth of the matter is that it isn't. The model is simply calculating probabilities based on the information it currently has.
Suppose that the beginning of a particular sentence is the following:
"The sky is . . ."
The AI language model might assign different probabilities to possible next tokens. For the purpose of illustration only, imagine that the AI language model has assigned the following probability percentages to these listed words:
"blue" = 75%
"cloudy" = 15%
"clear" = 8%
"banana" = 0.00001%
Please note that the previous numbers are only an imaginary example. Real AI models calculate probabilities over a very large vocabulary, and the actual numbers vary depending on the model and the exact context. The model then uses those probabilities to select a token. Once it has selected that token, the token becomes part of the growing response, or answer to the user's prompt -- meaning their question or comment.
The model then performs another prediction.
Then another.
Then another.
And another.
In simplified form, it would look like this if you could actually, and visually, see inside the language model:
"The sky is . . ."
↓
"The sky is blue . . ."
↓
"The sky is blue today . . ."
↓
"The sky is blue today, and . . ."
↓
and so on.
This process continues until the model determines that the response should end. Therefore, as we discussed earlier, the AI model does not normally retrieve an entire pre-written answer from a hidden database and simply paste it into the conversation. It is NOT an old school chatbot from the 1980s or the 1990s. It in fact generates the response step by step and TOKEN BY TOKEN. This is one of the most important things you need to understand regarding how a language model really produces text for its response output.
WHY DOES THE AI SOMETIMES GIVE DIFFERENT ANSWERS?
If the language model is calculating probabilities, there can sometimes be more than one reasonable possibility for what comes next. For example, consider the following:
"The little girl opened the . . ."
could be followed by . . .
"door"
"box"
"book"
"gift"
. . . or quite a few other possibilities for that matter.
Depending on the system's settings and the circumstances, the AI language model may select different possibilities on different occasions. This is one reason an AI can sometimes give different answers to very similar questions. As we have seen, the important point is that generating language is not simply a matter of looking up one fixed answer. It involves calculating possible continuations and selecting among them.
Please go to part three for the conclusion of this series.
⇒ Go To The Next Part . . .