How does predictive text work?

He went back to Markov's original idea of predicting text, but instead of just using vowels and consonants, he focused on individual letters. And he wondered, what if instead of looking at only the last letter as a predictor, I look at the last two? Well, with that, he got text that looked like this. Now, it doesn't make much sense, but there are some recognizable words like "whey", "of", and "the". But Shannon was convinced he could do better.

So next, instead of looking at letters, he wondered, what if I use entire words as predictors? That gave him sentences like this, "The head and in frontal attack on an English writer that the character of this point is therefore another method for the letters that the time of who ever told the problem for an unexpected." Now, clearly, this doesn't make any sense, but Shannon did notice that sequences of four words or so generally did make sense.

For instance, "attack on an English writer" kind of makes sense. So Shannon learned that you can make better and better predictions about what the next word is going to be by taking into account more and more of the previous words. It's kind of like what Gmail does when it predicts what you're going to type next. And this is no coincidence, the algorithms that make these predictions are based on Markov chains. - They're not necessarily using letters, you know, -Yeah they use what they call tokens, some of which are letters, some of which are words, marks of punctuation, whatever.

So it's a bigger set than just the alphabet. The game is simply, we have this string of tokens that, you know, might be 30 long, and we're asking what are the odds that the next token is this or this or this? - [Derek] But today's large language models don't treat all those tokens equally, because unlike simple Markov chains, they also use something called attention, which tells the model what to pay attention to.