Artificial Intelligence (AI)

This is a page about artificial intelligence (AI)

Fundamentals of AI

Model

Artificial Intelligence (AI) > Machine Learning (ML) > Deep Learning (DL) > Generative AI

This is the model from the AWS training course. It may seem a little weird to go from AI to Generative AI so I think the point is AI is made up of the following parts.

Training Data

It all starts with data – training data. AI is built on data and the relationship between data. The previous sentence is and contains data. The words are data and the whole sentence is a single piece of datum. This data is used to learn what the sentence “means” – the meaning, interpretation which requires a type of intelligence (to derive information from data can be seen as type of intelligence but let’s not dwell on what is intelligence).

Data can come in different classes, types and formats.

A class of data would be an image, video (moving image), text (typed and handwritten), sound – basically anything that a human can create.

A there are two types of data: 1) labelled – data that has meta or descriptive data cat = image and 2) unlabelled – data that has no label or description.

The other types of data are 1) structured – some relationship within the data such as in a table or time series or 2) unstructured – no relationship between the data such as text and images. It might seem a bit odd to classify documents as unstructured as the document will have a structure (sentences, headers, paragraphs etc) but the key difference is that one document is not related to other documents so all the documents together is unstructured (they may be the same topic but this topic not structured.

The Learning Process

Do you remember learning your mathematics times tables? You may have learnt by repeating the numbers in order or maybe through a pattern. Either way you had to take in information and learn from it and then recall if when needed. To get to a number you may have to go through the routine of finding the answer or just know it straight from memory.

With computers it’s a similar process. Data > Process > Memory > Recall. The computer has training data which goes to a learning process or more specifically a machine learning (ML) algorithm which then goes to memory or specifically a model. So it goes

Training Data -> ML algorithm -> Model

The data from before (labelled / unlabelled, structured / unstructured) goes into the ML algorithm which makes the model. When people use AI programmes you are asking the model to recall an answer to the question (sort of – the model can recall different answers each time but let’s move on).

There are three types of ‘learning’ (again learning is a just a computer program running rather than learning like a human). There three types of machine learning are:

  1. Supervised learning – this is for labelled data to create a mapping function for unlabelled data e.g. these are pictures of cats – cats got it – what are these? cats!
  2. Unsupervised learning – stage 2 – building a function from unlabelled data that is common across the labelled and unlabelled data.
  3. Reinforced learning – supervised learning with new data to and corrected to improve the patterns formed in the unsupervised learning. This is to tighten up the unsupervised learning with edge cases to test how the system performs.

Once the model has been built then it needs to tested and improved. This is done through inference. Inference is the model “inferring” or guessing based on what it knows what the answer could be.

Inference. Inference. They all got In-fer-ence!

There are two types of inference:

  1. Batch inference – taking in lots of data at the same time and then coming to a result based on the data and at the model (slow, detailed, complex)
  2. Real time inference – taking is small bits of data responding rapidly with a result that may be correct or need more information – chatBots!

Quick summary:

  1. Data is foundation of AI
  2. Data can be structured (tables) or unstructured (text)
  3. Data can be labelled (data about what is) or unlabelled (no metadata)
  4. Data (now training data) goes into the computer code aka Machine Learning algorithm and processes data to look for relationships
  5. Relationships between training data (and other sources) build the model
  6. The model is further trained (improved) through inference
  7. Inference can work in two ways
  8. Batch inference – lots of data to compare
  9. Real time interface – working with small amounts of data commonly through prompts

Going Deep – Deep Learning

Deep learning is the process of linking data in many ways. The many ways are seen as connections like you would see in a brain where individual neurons are connected to other individual neurons to make a neural network. Neutral networks in computer terms are functions (nodes) that build relationships within data that are connected to other functions (nodes) creating a super function between them with the strength of the relationship being a mathematical function in itself. These connections build up into layers. Layers create depth hence deep learning.

Deep learning is simply building relationships between the data. Some relationships are obvious – people in a certain area shop in shops in the area or young people like listening and going to new music. But there are relationships that can be hard to see or even impossible to see. Take for example a game of chess – a computer can see relationships between the pieces and it can store the relationships between these pieces based on the probability of future moves.

Two the common thinks deep learning can do are

  1. Computer vision – assessing a piece of data and returning what it looks like -“looks like a dog” (may not be a dog)
  2. Natural Language Processing (NLP) – computers that generate speak like responses to human input “I’m sorry, Dave. I’m afraid I can’t do that”

There are more but that will do for now as these are the things you are likely to either have heard of and actually use in everyday life.

Foundation Models

Data > Training > Model is the foundation of machine learning which we have had for a while. So, why now is AI such as big thing when it’s doing the same thing as machine learning? The answer is two fold: 1) more powerful processing and 2) more data aka the World Wide Web (www) to “train” off (for a little clarity the www are all the web pages, videos, images everything that you can get through a web browser (to be even more precise through the hypertext protocol (HTTP)).

The ability to read the WWW is not new. Search engines ‘robots’ first read the web page meta data (page name, search terms etc), then it read the contents of the page. This was the symbiotic relationship between the web site author wanting to be found and the search engine allowing people to find it along side some advertising. This was all done through conventional computing using central processing units (CPUs) found in Internet search engines computings.

To move from simply indexing the Internet to providing relationships between all the data and then generate new context require a new set of computer chips that were normally used to display graphics on a screen – graphic processing units (GPUs).

GPUs were perfect to allow calculations in parallel rather than in series. This piece of hardware along with some very smart maths allowed the contents of the Internet to be used to create very large machine learning programs that produced very powerful foundation models (FM).

Foundation models can be seen as the very well pre-trained models that other models can be built off (hence foundational). There are 6 steps to building using a FM.

  1. Raw materials – Data Collection
    Get hold of massive amounts of unlabelled data (like the World Wide Web) but it could be your businesses records. Whatever the source it needs to be big.
  2. Load into Engine – Pre-training
    Load the data into the first layer model builds a rough model. This creates a basic level of response based on the data analysed for example what are nouns (ball, house, person) and adjectives (blue, red, white) and from that some relationship (balls can be any colour, houses are white, and a person can be blue, red, and white (freezing, sunburn, race).
    The model can be trained on the same data over and over to get more and more relationships from the data. The aim is for the model to start to “understand” or more precisely to increase the probability that the relationship is correct – the green man is probably a pub (Green Man) rather a man who is green
  3. Refine, Reinforce – optimisation
    Improve the model even further with specific techniques. The amount you want to optimise will depend on the complexity of what you are trying to achieve and how much money and time you have
    1. Prompt engineering
    2. Retrieval Augmented Generation (RAG)
    3. Fine-tuning on specific data
  4. Stick or Twist: Evaluate
    After all the training you have to make a decision on if it’s worth using the model. Can the model be improved with more training – is it getting better? Is it ready to be released (is it safe or just stupid)? Or is it
  5. Release and Deloy
    Once you are happy the model you can make the model available to be used for a wider use of appliances. For example if you built a model called “Is it a cat?” you could make this available to the world so people can check if the picture they have is a cat. The model should be able to take a picture of a dog and respond with ‘Not a cat’. If it doesn’t then the model is rubbish and should be reevaluated.
  6. Continue to improve
    Track how well the model is performing (how many requests correctly answered, speed of response etc. Make some money or improve the world or maybe both.

Large Language Models (LLMs)

Now that we have our foundation models we can look a little deeper into the world that most people think about when we talk about AI – Large Language (foundation) Models or LLMs.

LLMs are exactly what that – they are a model that is based on a very large model that is trained on a language. That language does not have to be a spoken language like English of Mandarin but anything with structure be that a computer language, music scores or even images. Whatever the language the idea is to produce responses that are accurate in that language when prompted.

Transformer Architecture

The LLM is built using something called a transformer which is a way of taking the data (structured/unstructured, labelled/unlabelled) in breaking it up into bits (not binary bits) and then reconstructing the relationships in a way that questions can be asked of the data (either as text input or spoken) that can output a desired answer in a wanted format. e.g. produce a picture of a blue cat. The text has been transformed from text into an image.

Tokens

The input into an LLM is by tokens. Tokens are text inputs that the LLM uses to come up with an answer. For example “Produce a picture of a blue cat” could be separate into individual words on even parts of the word (compare understand with underwater – the transformer decides how to take the input and feed those in as tokens to the model.

Embeddings and Vectors

With inputs broken into tokens the tokens can be manipulated to see what possible relationships the tokens have to the model. The tokens blue and cat are given a 2 values or a vector to show where they fit in the model – when a token is given a weighting the token is embedded into the model. Our phrase blue cat is a very strange one. Blue is normally a colour but it could mean feeling down. Cat is normally a domestic cat so the model is looking for pictures of cats that are the colour blue which don’t exist so it will have to create one. The embedded token will be place in the colour and animal space. From there it can take a picture from its data and change a blue cat to blue. The model will deter which shade of blue and what type of cat. Image generation is commonly done through diffusion models.

Multimodal Models

LLMs can also take input that is text and images. This is multimodal as the LLM can take in multiple types of inputs and turn them into tokens. Tokens are then embedded and used to create an output.

Other models include Generative adversarial networks (GANs) and Variational Autoencoders (VAEs)

Optimising Models

Prompt Engineering – instructions to the foundation model. PE is how to make the input layout, process and instructions easy to use. Commonly the prompt may ask several questions to understand what is being asked.

For example “Create me a picture of a blue cat”

“Sure. Would a pet cat that has blue fur ok?”

“Yes”

Generate blue cat.

“That is not what I am after”

“How you like me the change it?”

“Make it a lighter blue”

Whatever the feedback the model is a way of optimising.

AWS Infrastructure and Technologies

AWS has been at the heart of the AI transformation due to the extensive Cloud infrastructure and on that the ML and AI tools. AWS breaks the ML/AI tool into three layers

1 Amazon SageMaker AI – build, train and deploy ML models in one place

2 AI/ML Services – 7 classes of service – these are ready to use

  1. Text and Documents (reading)
    1. Amazon Comprehend
    2. Amazon Translate
    3. Amazon Textract
  2. Chatbots (prompting)
    1. Amazon Lex
  3. Speech (hearing)
    1. Amazon Polly
    2. Amazon Transcribe
  4. Vision
    1. Amazon Rekognition
  5. Search
    1. Amazon Kendra
  6. Recommendations
    1. Amazon Personalise
  7. Misc
    1. Amazon DeepRacer

3 Generative AI

  1. Amazon SageMaker JumpStart
  2. Amazon Bedrock – access to pre-trained models and APIs
  3. Amazon Q – generate code
  4. Amazon Q Developer