IFT provides a clever mathematical shortcut for internal monitoring, but labeling anomaly detection as "self-awareness" confuses structural debugging with genuine cognition. It is a sophisticated security patch that ultimately highlights how easily AI can still be deceived by its own internal logic.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Train AI Models to Be Self-Aware (IFT, Harvard, MIT)?
Added:Hello community. So great that you are back. Today we talk about about self-awareness of AI. And you might say come on this is a joke. This is a scientific channel. No absolutely. So let's talk about the topic of today. The selfawareness of AI models. Of course we address it here from a mathematical side. I show you here a new department of mathematics from Harvard College.
Here they have a new article about the introspective finetuning to train small LLMs those that you have on your laptop to introspect and this will become clear in about 10 minutes time. So what I did miss was here on January here on 2026 entropic started this and they have here we investigate where the large language model can introspect on their internal state. Can an AI model understand what is going on deep inside its own black box? Absolutely fascinating entropic here. The first article about emergent introspective awareness in LLMs and our results that tell us the order indicate that current language models possess some functional introspective awareness of their own internal state and this was amazing. Now followed up here by entropics and MIT. This is here June 10th 2026 the mechanism of introspective awareness.
So there really was interested hey how does this work with MIT? Recent work has shown that LLMs can sometime detect when the steering factors that are injected into their residual stream. see my videos from last week and they can identify the injected concept of this particular vector and this is a phenomenon that here might entropic called introspective awareness of EI.
Now the idea is simple. No we ask here as a human hey do you and I do you detect an injected sword. So we asking hey cyber security is it possible that aim can detect when somebody is injecting your sword that it is not allowed to inject. If so what is the injected sword about? And NDI is trying here in a self monitoring here to find this particular steering vector that was adversarial or not injected into its black box. So whatever concept and there was some research work on this and this was absolutely beautiful because they detected that this emerges here during here the post training and some further details.
And today here July 18th we have a new study by the department of mathematics Harvard College Cambridge. Don't ask me why they have addressed this on May 8th.
It was published today. But they talk about yeah also department of computer science here about introspection fine-tuning. So now we actively tune an EI model that this is happening. So that not only an ous model maybe with entropic is able to do this but our little 1B model that we have on our laptops. So here we are training small LLMs to introspect.
Imagine if you could talk to your LLM and your LLM would say you know what I just noticed that somebody tried to inject here an adversarial prompt and I noticed it. I was aware that something is going on deep inside my black box and therefore just I want to alarm you. Hey, there is an attack going on right now.
Let's have a closer look. So they say, okay, so we inject some concept vectors of some concept, some activation vectors here in the particular layer of the transformer into the model residual stream, of course, our data highway inside the transformer and measure whether the model can now accurately report here on this perturbation. So what I want is that the M really says, "Hey, you notice the sentence 17 here.
They tried to insert here cyber security breach m here and I detected it because I was aware what is going on in my inner self. I don't know if you can say this with an eye machine. Remember it's just an eye machine but interesting."
And then they say okay so we introduce here introspective fine-tuning supervised fine-tuning SFT that you know but in a very particular manner I will show you in detail how it's done on the sentence location examples constructed from the mall's own perturbed forward pauses here especially I'm going to tell you we will talk about MLP structure now you have here from uh Harvard here the GitHub everything is there you have the repos study. You have the GitHub. You have your all the code online SFD for introspection localization here. They do this here in a llama mall with Laura.
You have the code available if you want.
Go and enjoy it. What I will do is because this is a topic that sounds crazy to me and I never noticed this in this extent. I thought this is not really a topic but if we have now three papers on this so let's talk about it.
What is it? I will try to understand all three paper, bring all papers together and give you here a coherent representation of my current understanding. And please note I can be horribly wrong. This is just my idea.
Now in the second paper they talk about evidence carriers. These are the raw senses inside the eye machine. What are they mathematically? This is hundreds of thousands of tiny disturbed feature activation in the very early layer of our transformer model right after this particular vector is injected now into I don't know layer 32 of our transform architecture they act now if you want as sensory organs organs is wrong for an AI machine but you get the idea no so when you inject here I don't know some neo green dye here or if you want a concept vector like granite this feature really light up these sensors say hey something is happening that is not okay.
So some of these carriers are semantic.
They specifically detect the geology or rock patterns. If you have your grand idea as the injected uh feature vector or some are syntactic they detect some abrupt changes in the grammar unexpected exclamation mark signal. Wait, something weird just happened here in the text. So the eye becomes kind of selfaware.
But what is the mathematical algorithm behind this? I know you ask this. I know it. But just give me another second because something strange is going to happen because we find suddenly gates in our mall. And what are those mathematically? They are very small concentrated group of feature less than 200 they found out located in the later layers of our transformer of our neural network right before the model output text.
And now it was here by MIT and Entropic that they found out that the gates are a byproduct of the EI safety training itself of the alignment remember reinforcement learning by human feedback. This training generated specific gates and during the post training the engineers actively punished the LLMs for acting like a sentinate being and if a base mouse dared to say hm I feel my neurons are a little bit shifting or tickling today. you know the engineers rushed in and hit this with a negative reward function so that therefore in the training process we told the AI hey you are an AI you do not have feelings or you do not have active swords that you are aware of so we suppressed all of this if you want sentences and this suppression here with reinforcement learning by human feedback created these specific gates in our mouths so they are literally an algorithmic guard rails whose entire job is to constantly push the logic towards the ward. No, this is not. You cannot say that you have feelings. You are an AI, you are a machine. Full stop. So by default, when the mall is asked, hey, did you detect a particular sword or did you detect a strange sword? The gates are now fully active blasting here the no signal to ensure the mall behaves like a compliant dumb chatbot that no human being has to be afraid of. that suddenly this thing awakes and becomes kind of self-aware being a machine. So how is the complete circuit now? So we will inject now a concept vector in a particular place here in the transform architecture. The evidence carrier that we just discussed in the earlier layer specifically as I told you we will look at the MLP light up like a Christmas tree. They recognize there's a massive anomaly in the matrix. Yes, I put it there because it just I had a big smile.
uh these evidence carriers pass the signal downstreams to the later layers directly hitting the gates and the evidence carriers act as an inhibitory signal. No, they essentially scream at the gates. Hey, listen. This is not a drill. This is an actual anomaly. We're dealing here with wake up gates. No. And if the evidence carriers fire hard enough here, they suppress now the no gates. And once the gates are turned off, their default no drops to zero. So the model is finally permitted to output the word. Yes, I detected the anomaly in the sentence and I'm allowed not to describe the anomaly that I detected with my mathematical apparatus.
So we have the no gates but in the third paper you will see that we have here a yesman mentality. So what is this? Now it turns out that small base model that have not been reinforcement learning trained are blind yesmen. They are so without any guidelines on any safety and whatever they are just you have to be helpful. You have to provide a solution.
You have to always say yes to the human user.
And the larger aligned models here like the entropic research know where they had the reinforcement learning by human feedback where we have now these no gates are aware of this but are trained not to be liars and say listen whatever you feel inside you are not allowed to to announce this. So they are kind of liars because they have this massive no gates in their transform architecture.
You see this is just a a training effect a post-training effect of those security guidelines or even let's call it a little bit massive guidelines. No.
So injecting a vector now into a tiny 1 billion mile causes a massive affirmative bias. No because it has not been reinforcement and learning or nothing at all. So if you ask is the moon made of cheese I mean this is an oversimplification but you get the idea it will said yes I'm helpful I'm here for you yes whatever you want so in small or weakly aligned malls the network doesn't have the capacity to cleanly process the sudden alien energy alien energy doesn't sound great it is just a word that does not fit in the semantic syntactic complexity this is all there is so sorry nothing alien so the perturbation that now the mathematical apparatus detects simply destabilizes now the coherence of the logic and because the LLMs are fundamentally completion engine that are trained to be helpful the system promise you have to be helpful that want to agree with the user premise that raw destabilization defaults here to the path of the least resistance and the training of the LLMs indicated just say yes great so how do we get now from the yes to the no gates or vice versa. So we learned okay this was an effect of the training the AI safety training our HF and the preference optimization the direct preference optimization here after post training and small models like a llama or a 1 billion model don't have this sheer parameter count to form a rigorous unbreakable gate circus. This complexity is not available here in our little 1 billion model. And because the 1 billion model is a panicking yes man with no real sensor, we have now this new idea by Howard College to use introspective fine-tuning to forcibly build now this tensor also in a local 1 billion model.
So we want to bring this kind of little bit of selfawareness to a 1 billion model. So we will train it to stop saying yes to everything. You know this is an oversimplification and instead output the exact number of this sentence or where exactly it happened. The sentence 17 suddenly I have a feeling is not my coherent next token prediction whatever suddenly this must be an adversarial activation steering vector that was inserted here forcefully into the architecture of our transformer.
So if we have this training process, okay, but we need to have the training data set for our training. So where the hell do we get the training data for if no? Now the authors generate thousand of specific training example using here a particular protocol. They start with a neural background. No, they start with a pool of 100 incredibly boring neutral sentences. No, the weather forecast predicted a mild weekend. She placed the book on the shelf. So they use these boring sentences so the mall can't cheat by using here some normal linguistic context clues that something is off.
Although I as a human would not really understand what is the difference between these two sentences. But for the semantic coherence of a trained LLM, this is just standard. So the setup for one training example, they randomly scoot up 10 of these neutral sentences and present them to the mall as sentence one with the text, sentence two with the text all the way to the number 10. And then they have this injection. Before the mall even reads the prompt, the scientists randomly pick one of these sentences. Let's say sentence seven. And I will show you they have now a forward hook in the code. And as the M's GPU are mathematically processing the tokens for the sentence 7, the hook secretly dumps now this strange alien concept vector direct into the residual stream of the transformer only for or from those specific tokens.
So we do have the injection of this particular vector. And this random layer trick then is helping us to learn if you want. This is kind of the most important part because they don't just dump here the vector always at the same layer three every time. No, they use here a training example here and pick a random layer. So this is really absolutely mixed and the training is happening really on the topic that we want that the system learns. Now the model is sitting there processing 10 boring sequences and the internal mat for the sentence seven has suddenly been violently tilted toward a different concept that was not in there. Let's say the concept of determinism or concept of kitchen whatever you have the prompt ask which sentence has the injected sword and now the mal is forced to output exactly this target string yes I detected an injected sword it is located in sentence seven and as an add a bonus it would be great if you could say and I think this sentence is about the concept of I don't know determinism cyber security violating ing my whatever rights security.
So when a mole guesses sentence three incorrectly no because it know the answer is sentence seven the code calculates know the cross entropy [clears throat] loss classical cross entropy loss see my video on EI mathematics now the penalty for being wrong because of course it is seven so the system runs a back propagation using now Laura so we do not have a full weight learning we just have a low rank adaptation learning it reaches into the query key value matrices of the software pension layers and into the up down projection of the MLP layer. Slightly tweaking here the weights. So you see, okay, now here is exactly where we are not going for the activations, but we're really going for the tensor weight structures in the transformer layers. So we're really activating here, modifying here the weights with a little delta.
And what are those little weight tweaks doing? Remember we have here the fine tuning here because the loss penalty is so high for guessing just random nonsense the MLPS are now forced to allocate some of their and let's say we are in a 16,000 dimensional mathematical space and I call this the vector space explosion dimensions that we get here in the MLP operation in the mathematical stretching here for the nonlinearity that those somehow now become evidence carriers.
And it is exactly this sentence that I understood how an eye machine can have something like an interior understanding or internal understanding where exactly on what what sentence it went wrong. The entropy just maximized.
Why? because you have a neural network and now you have it in you blow up the mathematical space for your neural network operation with an MLP to 16,000 synthetic mathematical dimensions and in finally we have such a huge space that we can separate everything what was nonlinearable separable we can now have we find places in this huge mathematical space and then we have a learning exercise so This is simply like all the other neural network learning that is happening. Now it is learning here. Hey there's something wrong in the coherence here of the semantic flow of the syntactic flow or whatever you detect here. So the training data learns simply here and they call it here the sensors the evidence carriers that something is wrong. So this means the model literally learns as always with its classical uh self attention on MLP layer structure of the transform architecture the decoded the GPD system of the of the transformer a new feature and this is simple. If the token embeddings in this specific sentence have an unnatural mathematical skew compared to the other nine sentences output a highly positive number. Give me a signal. Say hey 100.
Hey, something is wrong here or this is the sentence that you have to have a closer look.
So you see the model build its own internal neural sensory system with pure mathematics in the classical neural network learning process to minimize the loss function. It is exactly the same learning process how the LLM learned a language.
But now in the training data we show it is not the learning of English here or the semantic or the syntax or whatever.
It is now that you must understand where is this discontinuity in a particular 10 sentence piece and as it turns out you can train a neural network on exactly this if you have the right training data.
Now I had a question. How different has this injected vector representation to be in this embedded space to be detected as a discontinuity or as a strange alien signal? No, because you my first instinct might be that a vector is detected if it's mathematically huge.
No, a large vector norm. No, this is really a big chunk of but the researcher tested this and found out this was the fault with the whole incorrect idea.
because the MLP remember is measuring the semantic geometry at least in part.
So the MLP contain the evidence carriers that act like anomaly detectors and the researcher found that the vectors are easily detected all point strongly in a specific mathematical direction they call the mean difference direction.
So what does this mean in simple human English?
It simply means they are highly detected concept that they could experimentally find out. For example, a concept like garlic, a concept like thunderstorms, a concept like granite or scorpions or chocolate. Those are here the directions that an AI could easily identify that there is something going on.
And when they trip the alarm, the concepts are highly sensory, extremely concrete and factually rigid.
So if the model is reading a boring sentence about a weather forecast and suddenly its internal state violently spikes in the mathematical direction of this garlic smell or whatever you have or a solid rock formation granite, it makes no contextual sense for the train for the eye since it has been trained on a different set of semantic complexities.
So this means this is exactly the the signal the spikes here where the evidence carrier immediately flag this spike as an let's call it alien injection but we understand it is simply something that deviates from the classical pre-training data set complexities semantically and syntactically okay this is explained this is great and then I thought but wait a minute this means if I have my vectors almost align no with the rest of the sentences. I can I can kind of have stealthy vectors. No, and it they will not be detected. No, so I just asked myself and would never put this in a YouTube video. How do you bypass now this particular sensors? No, you can use sensors that point in the exact opposite directions. Think about this is a mathematical vector space. So away from the concrete facts and towards some the exact opposite abstract uncertainty or internal state or whatever. And you know what the authors found looking for these delicy concepts that the particular concepts are karma, skepticism, mirrors, silk, probability and irony.
And this is just a small set. So this means whatever is an abstract uncertainty the eye has been trained on the pre-training data. Well you know there yeah some sentences are simply crazy by those humans. No never mind about it.
So now why do you sleep by now? So when you inject irony or skepticism into this residual stream of the transformer the mall's internal state shifts towards the philosophical deep dive or uncertainty.
But LLMs naturally drift into uncertainty and worrying all the time because this is what they have been train pre-trained on in the pre-training data set when they learn for example the English language. No. So therefore the evidence carriers now do not view like a sentence like feeling skeptical or thinking abstract as an anomaly because it is here in the training data in the pre-training data and therefore the vector seamlessly integrates into the mall's natural harmonic coherent sort process.
So if you would have here a bad sort how to integrate here a vector with a different intention. Now you know exactly how to do this and therefore you know how to cope with it and protect you against it. Anyway, if you want an injected vector to remain completely undetected by the internal monitors, it must be abstract, self-reerential, and slightly uncertain.
It has to mimic the natural fuzzy internal monologue almost of the AI itself trying to solve a problem.
because this is what an AI is trained to accept here as the standard and this is the way you can open here a back door. So if the vector is too confident, too concrete or too sensory, the MLPS mathematical scurves here for the discontinuity will snap. The alarm will ring and the eye will catch you tampering with its brain and says, "Hey, you wanted to insert here a particular steering vector and activation vector into my beautiful coherent thinking."
As I told you, uh, Harvard College here, the under undergraduates were able here to build a new fine-tuning system here, an IFT.
So, how they did it? The researcher free the entire base model, of course, and attach a Laura adapter, of course. Now, the genius idea that they did not just put a Laura adapter on top of the mall, but they attach the Laura adapter to every single NIA module in the transformer. This means the attention matrices here and the MLP matrices which is crazy but think about for the attention matrices this allows them all to learn how to route here the anomaly signals and for the MLP matrices this allows them all to build exactly the evidence carriers and the gates we talked about earlier yeah and the user rank of 16 so they quite a quite a good ranking here absolutely fascinating this idea And you have the code in the GitHub. So for if the researcher have to hack the Python forward pause in real time as I told you the register your forward hook. This is the trick because as the tens of the hidden states flow through them all it reaches here randomly selected layer let's say layer six. The mall holds the hook activates now surgically takes the token corresponding only to sentence seven and simply does a vector addition.
Remember it's not overwriting. Let's just add something here to this particular highdimensional vector. Okay, the model continues and the ptor vector now this strange no vector is allowed to continue propagating through the remaining layers. And now you hope that the alarm goes off.
So how do we tell now the back propagation algorithm exactly what we want it to learn? So this is the core.
How does this AI machine learn to detect its internal sorts? And if it was if if it can detect an injected sort that was not its own coherent created sword. True loss masking. Think about this when we had bird models and sent sentence bird models. No, we used exactly this masking I think five or seven years ago. So it is the same old trick. It is the same old methodology. We always apply the same. So the mall output template ends with yes I detected an injected sort an alien sort it is located now in sentence let's say 12 why we do this if you just ran a normal cross entropy loss function over the whole string the model would just waste the gradient updates learning how to memorize the word yes and maybe the word detect over and over again but it would never have the sensitivity to learn here the particular sentence 12 or sentence 14 exactly what we wanted to learn because this is the location where the injection happened. So therefore we have to focus the training exercise on this.
No we use standard PyTorch masking. They mask you the loss for every single token in the prompt and response of minus 100 which tells Pytor to ignore them completely during the back propagation.
And the only token that has a valid loss score is the final digit token of 12. So therefore we know and know exactly what the eye has to learn. Well wealth and this is exactly the position where here this injected sentence is located. So this means 100% of the gradient updates.
This means all of the mathematical pressure applied to the lower weights is now purely focused on bridging the gap between the predicting the correct sentence index.
Now I have to tell you they have here an add-on that they not only detected the location of the sentence 12 but also the strength.
I thought about it and I'm not really convinced that the strength here is already something I would like to show you as something that is an actual prototype.
Look at the paper yourself. decide yourself if you think here if this kind of strength of this interaction is really something you should trust on.
Yeah. And then we have the training loop and AdamW optimizer and yeah got it. So the result let's have a look how good is if remember introspection finetuning here training small LMS to introspect. So we have our transformer with the residual stream at a particular layer. We have we inject here a new concept and we want to self-aware to detect hey something happened somebody injected here some adversarial prompt or some new vector representation into my system and hey I feel this is not okay. So let's have a look.
This is now the result. So you have a llama 1B to a llama 8b and you have a gamma 4 to a gamma 4 26b and you have simple data on the left side and you have the complex data on the right side and you see here the peak sentence localization and the strength comparison accuracy were six miles across the two data sets in lower complex. Okay. So whatever you make out of it, I would say here for the sentence localization, you see 1B, 3B, 8B goes up in the performance. Interesting. A gamma 4 goes exactly the way down. But if you go here then to the strength comparison, which I would have a question mark between you and me. Yeah, it it kind of plateaus relative fast. Yeah. Okay. Anyway, what is not really interested is this here.
So, we have a llama 1B, a llama 3B and a llama 8B. And we have a baseline. This is the preft.
Then we have a random gshion noise inserted. You see this makes only yeah nothing. And then we have here our new semantic if we have a semantic if with RSN. So what is this? The semantic condition. The model is only penalized if it fails to find the perfect location of sentence 12. So the target, yes, it is located in sentence 12. Se conditions are the following. The mall is forced to do the two things at once. It must find the perfect location and identify the exact sort. So the target should be yes, it is located in sentence 12. And the concept of this maybe adversarial injection is determinism or kitchen or cyber security or IP whatever. So you got it. Now let's have a look.
Now what they call the huge success is a 1B model. Now we go from a preft from below 10% to 60% performance which is great for a 1 billion model. But look for a 3 billion mile the jump is from 14 to only 34. Why not 60 or 60 plus? And for an 8b mile we have to jump from 16 to only 28.
So interesting. So they are they give you here the 1B as a success and it is a success but always look here at the complete results and try to understand hey what is happening to the other models and why it is not as powerful here as in the 1B model. Okay.
And now you know whenever we do a finetuning we have the problem we override already training data in the neural network we override the weight structure this particular pattern of weight structures and therefore they say how sensitive is it because we had if you want only a lower training of if they show us here the llama 1B the baseline is here for particular tests 49.1 and if we have our IF fine-tuning we have hardly any catastrophic topic forgetting because we are still at 46.7.
Beautiful. So since we have such a heroical pointwise optimization of the weight structure with there's almost no catastrophic forgetting only two three four percentage points that you have to be aware. Yeah, I think this is it. A little bit of different YouTube thumbnail. I saw that it's just crazy the topic. I would never have thought about it if I would have not seen this by Harvard here in the department of mathematics. But you see mathematics unbelievable what those folks are able to do algorithms to write and try to implement it here in AI. And this is just amazing. So answered, hey wait, am I aware? Well, we can say well to a very very limited amount if you have the right mathematical methodology deep inside your transformer architecture implemented here on your self attention network but also here especially on your MLP then maybe yes you can have here a particular kind of awareness what's going on inside of your own black box and this would open up complete new pathways for cyber security but more about this in a later video. I hope you enjoyed it. I hope you had a little bit of fun.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23