This research presents attention self-modeling as a promising methodological approach for studying consciousness in AI systems, distinguishing between phenomenal consciousness (subjective experience) and access consciousness (information available for reasoning and control). The team argues that attention is unusually influenceable and inspectable in AI systems, making it a tractable target for theory-informed testing. Drawing on Attention Schema Theory, which proposes that consciousness depends on an internal model of attention, the researchers suggest that AI systems with attention self-modeling capabilities may exhibit behaviors associated with consciousness. Empirical evidence from multi-agent reinforcement learning and large language models shows that attention self-modeling supports prosocial behavior, task categorization, and emergent attentional mechanisms, providing a bridge between abstract consciousness theories and measurable, manipulable features in artificial systems.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Frontiers of Self-Attention and Artificial Consciousness - Angie Normandale and Sahba Afsharnia
Added:Hello, can you hear me?
Yes, fantastic. Great. So, welcome to our paper presentation on frontiers of attention and attention self-modeling and artificial consciousness.
So, uh next slide, please, thank you. We will do a quick introduction. Who are we? Why are we here? What are we doing?
We're going to talk about the links between attention and consciousness um in terms of theory. Then we're going to talk about some evidence from different kinds of current AI models, including multi-agent reinforcement learning and social cognition, and also self-monitoring in large language models.
Um finally, implications and discussion.
Cool. So, why now? Why this?
Um who here has read Anil Seth's Bergmann essay? Yeah? Anyone else? Yeah?
I know you have cuz we've talked about it. Um Yeah, I've read it, too. Um this is a very interesting essay for Bergmann Institute, who are groundbreaking, to choose to win the Bergmann prize. Anil argues that consciousness can only exist in biology.
One of his main arguments is that biological agents do not compute, and therefore computing is just fundamentally different.
Um this remains to be proven, unfortunately.
He He also argues that simulating something is just not equivalent to the same thing. So, we can build something that will simulate different behavior, different mechanisms, it's not the same.
And he says it's not the same because it won't have the same causes, and it won't have the same results.
Well, I have an answer.
Prove it.
This sounds to me like a challenge, and this is how our team saw this, and we we that we would go for empirical evidence in the spirit of, you know, those first days of neuroscience where we've got a thing we don't know how it works and we're going to find out.
Um, another frontier paper that that reminded me why I wanted to do this talk today. Has anyone read Eric Schwitzgebel's new book?
Anyone? No, you should all read it. It's really good. He says that we are all absolutely ruined. We are stuck in the fog and we're never going to figure it out when it comes to consciousness.
He says, "That's it. It's too late.
We're going to make hundreds, thousands, millions of conscious beings. There's nothing we can do. We won't know they're conscious. We won't know they're not."
Um, and his argument, he goes through different theories of consciousness, explains exactly why he thinks we're stuck in the fog. And his reasoning is that there are all these different aspects of these theories that are just deeply underspecified.
There are no minimum conditions, no boundary conditions. What are the actual mechanisms? What are the results? And how do we know what matters?
Again, I see this not as a reason to give up, but a challenge.
And I appreciate other people like you in the room doing the same.
So, who are we? My wonderful, wonderful Enveloteam, um, two of us are here today. We have Dr. Joel Peuker, who is a Finnish He's our lead machine learning researcher. Rasmus Halo, he is a theoretical neuroscientist, um, and Mack and Turper.
And Sabah, who's come from Toronto today. And me. Um, I am half a computer scientist, um, half a neuroscientist, and then a little bit of other.
Cool. So, we have some empirical research going on at the moment. Um, we don't have any results because we submitted the paper very soon into the start of our research. So, we decided to write this position paper explaining why we're doing what we're doing and not something else.
And what we think other people should do to further the field.
So, that's what you learn today.
Okay.
Cool. I think it is time to hand over to Saba, who's going to talk about the theoretical background to kind our work.
You could also speak in here.
>> Okay.
Um so, hi. Um to start um the question of whether machines could be conscious is no longer just a science fiction question or just like a philosophical debate.
It's increasingly becoming a practical problem about how to reason under uncertainty.
So, um even if we don't know whether an artificial system is conscious or not, we still need to ask what kind of evidence should make us raise or lower our credence that it might be conscious and also what ethical or governance responsibilities would follow if that uncertainty can't be fully resolved.
So, this is the kind of issue that is raised by the Butlin et al. paper in 2023.
At the moment though, both public and scientific discussions are still often being pulled into um towards two unhelpful extremes. So, on one side, there's a tendency to rely too much on self-report or anthropomorphic intuition or fluent language output. And um this is where a system says uh something that sounds introspective and people take that as meaningful evidence. On the other side, we have the opposite extreme where people dismiss the issue entirely with broad claim like AI is like a brain or AI is nothing like a brain. And neither of those positions really help us move forward scientifically.
The main methodological problem here is the measurement problem. Even if we start from a serious theory of consciousness, it's still very difficult to turn that theory into falsifiable tests that could actually be applied to artificial systems. And without falsifiable tests, we end up doing post hoc interpretation, where we look at a system after the fact, notice some interesting behavior, and then retrofit a story onto it instead of making real empirical progress.
Uh so, the central argument of our position paper is that attention self-modeling is a promising target for tractable theory-informed testing.
And I want to be clear about what that claim is and what that claim is not. So, it's not that attention self-modeling will solve the hard problem of consciousness, and it's also not the claim that attention self-modeling is sufficient for phenomena phenomenology or subjective experience.
But instead, we're making a narrower and more methodological claim that attention is unusually influenceable and inspectable in AI systems. And it also plays a central functional role across multiple theories that connect attention to access, control, and reportability.
This matters because it gives us a very plausible bridge between theory, measurement, and implementation.
And instead of staying at the level of very abstract claims about consciousness, we can focus on a feature that's already discussed in several major theories, and then is at least measurable and manipulable manipulable in AI.
Um because the word consciousness is often used in a very ambiguous way, uh we also try to specify what exactly is being tested. And uh a standard distinction in literature is a distinction between phenomenal consciousness and access consciousness.
So, phenomenal consciousness is the what it is like aspect, like for example, how Nagel says it. And access consciousness, by contrast, is the information that is available for reasoning or rational control of action and for reporting.
The distinction here matters to us because the relationship between attention and these two forms of consciousness is not symmetrical. And phenomenal awareness sometimes can happen with little or minimal top-down attention. But access consciousness is more tight more tight to attention-mediated selection and prioritization and transmission.
So, the focus here in the paper is mainly on access-related and report-related and also self-monitoring capacities.
And we treat phenomenology as an open question because of the metaphysical limitations.
It's not denied, but it's also explicitly bracketed because any serious test has to either operationalize it or very careful operationalize it in a very careful manner or admit that it remains unresolved.
So, overall, instead of asking in a very vague way whether AI is conscious, we ask a more tractable question.
And based on the premises of these theories of consciousness, can we identify and test specific properties related to attention, self-monitoring, and internal control that would count as theory-relevant indicators?
So, the next question is why focus specifically on attention self-modeling?
One reason is that current frontier AI models increasingly make claims about being conscious, and those claims are creating real confusion.
Landemore notes this directly.
And even more interestingly, Berg, Delucena, and Rosenblatt report that training a model to be more honest can actually increase the frequency of these consciousness-related claims.
And that raises an important question, that when a model describes an internal state, what exactly is it referring to?
Is it simply producing fluent pattern completions based on language on its training data?
Or is it referring to stable internal variables that the system genuinely uses to monitor and regulate its own information processing?
And if such internal states do exist, then to what extent are they analogous to anything we would associate with conscious processing or even valence consciousness?
And attention is a strong place to investigate these questions because it can be studied at multiple levels. At the behavioral level, it can be measured through reaction times, error patterns, or lapses in sustained attention. At the neural level, it can be linked to shifts between large-scale brain networks like the default mode networks of control control-related networks.
And in AI, it can be studied computationally through inspectable and um intervenable mechanisms at the level of architecture, representation, and processing dynamics.
Um so, attention is also especially useful because it plays a functional role across several leading theories of consciousness.
In global workspace theory, for example, attention functions as a bottleneck that helps determine what information becomes globally available for control, reasoning, and report. In higher-order approaches, also, consciousness is linked to metacognitive monitoring, which often involves attention to one's own mental state. And in attention schema theory, this emphasis this emphasis goes even further where uh consciousness is tied specifically to an internal model of attention itself.
Uh so, the reason to focus on attention self-modeling is that it gives us a tractable and theory-relevant target.
And it's something that we can potentially inspect and test, while also comparing predictions across major theories of consciousness.
So, what exactly is an attention schema?
In Michael Graziano's work, attention schema is an abstract and simplified model of the most basic properties of attention. That includes its dynamics, its consequences, and its constantly changing state. So, it's not attention itself, but a representation that a system uses to track, predict, and regulate how its attention is operating.
And this is important because the ASC proposes that conscious awareness depends on the contents of this internal model.
And one way to think about this is by the analogy with the body schema. A body schema is not exactly the body itself, it's a model that the brain uses to keep track of the body in the way that it supports its movement, coordination, and interaction with the world. And in the same way, an attention schema would be a model of attention that supports the regulation of its cognitive processing.
And Graziano argues that this kind of mechanism could help explain why a highly integrated and predictive machine might display behaviors associated with consciousness. And in that sense, the attention schema is interesting because it directly links architecture and behavior. It's not just a vague philosophical proposal, but a potentially measurable and implementable feature.
If it's really uh an important gatekeeper of consciousness, then if we investigate it, we could eventually help uh in predicting, testing, and perhaps even controlling whether an uh artificial system should count as conscious in some functionally relevant sense.
And once we accept that the next question is more operational. What would a model of attention actually look like, especially in a way that could guide AI research?
So, AST says that the brain constructs a model of attention and that the conscious experience depends on the contents of that model. But to make this useful for you for AI, we need to move beyond these familiar metaphors and especially the idea of attention as a spotlight. Even though the spotlight metaphor is very intuitive, it's too static and too simplistic.
Um Okay.
We will Do you want to >> Yeah. Yeah.
Okay, how about um Come on.
Why don't we summarize?
>> Yeah.
>> I'll summarize this bit and then we'll be going to the AI stuff. Yeah.
>> Okay.
>> Yeah.
Thanks.
Thank you, Saba. I really appreciate it.
>> What's >> No, it's okay.
>> Okay.
>> Um yeah, so I brought Saba on board the team precisely because we needed somebody who is a theory expert to actually try and make sense of Graziano's work because he doesn't specify what an attention schema would look like. He just says, "It's like a body schema but for attention."
And this was not enough. And I really enjoyed this diagram from Watzl because suddenly we have an ordered list. So, "Oh, I can find a list. I can look for a list. I know what a list is. I can do dynamics on a list. I've done data science." And this suddenly became quite exciting to me. Um the idea of searching for something in AI.
Right, let's move on and briefly talk about some stuff that we've got in Frontier AI animals.
Okay. So, we're going to be really brief on this.
Um there have been very few, just starting to be, studies about attention schemas in AI. In multi-agent reinforcement learning contexts, we see an increase in prosocial behavior, including agents becoming more interpretable to other agents when they have a model of their own attention.
They're also better at collaborating.
And uh yeah, and the reason for this is that they get better at task categorization.
So, all this evidence very very slowly, just a few studies, starts to support the idea that in AI systems, fundamentally, your model of your own attention is not just intelligence, it also supports this social component, which is something fundamental about consciousness.
Now, in large language models, we're going to skip this. Um yes. Okay, we're going to skip this, too.
Cool. Um this gets quite interesting. We have this idea of introspection, and when we look at some frontier studies of introspection, the goal seems to be to find out if your model knows it's being evaluated and is going to go scheming on you. That's the canary string.
And this is a fundamentally different motivation to studying metacognition in people, and funnily enough, the studies look very different.
Um okay, so, very briefly, because uh Keenan's going to talk about some of these anyway. Um in fact, I'll let Keenan talk about it, but there is some very new evidence about model introspection. It's very basic, there's a lot of flaws.
Cool. Um one of the biggest and most interesting flaws of this work is that we find that there are just different tests being employed to test metacognition in language models and in people. So, the literature on access consciousness from neuroscience is about what we call pre-cognitive and post-cognitive judgments. I say to you, "Hey, how much of this talk do you think you're going to remember? Would you bet that you'll remember like all of it or like half of it?" That's a pre-cognitive judgment. And then, I make you do a test.
And then I test then I say, "Okay, how much like did you think you remembered?"
And this is a post-cognitive judgment.
Um and then I check your answers against the test to see how good you are at making cognitive judgments, which gives me some kind of ground truth. And it turns out that people uh people's skills at metacognition are independent of their IQ. It's not just that you're intelligent and good at making judgments about yourself. There's like an extra seems system to be able to do this. And this is interesting. But, it turns out that the LLM research is asking, "What do you know?" not, "What do you know about what you know?" And that might be quite important to ideas of first, second, and third-order cognition. So, in one of these key Anthropic papers, if we go back, um we find that Jack Lindsey at Anthropic claims that these models able to learn something about themselves, to tell you something about themselves, is evidence of access consciousness.
But, it's sounds to me like it's evidence of memory.
Um and I'll leave that to others to discuss.
All right, okay. Final bits. Um so, some really interesting things that we have very recently found through mechanistic interpretability. There are some emergent, not trained in, emergent mechanisms for attention.
And we see things like um attention heads in one layer providing information or context for other layers, and specializing over training. Um so, things like inhibiting some competing context to make sure that you're in the right context further down.
Um we see tagging information to make sure that it's bound to the right bit of the other information in your input. So, things like, "I know that sub belongs to antelope." and I'll tag those two things, and then continue my processing knowing that that's who that is.
Um we see function vectors, this kind of context window idea, and a few other things that I'll let you investigate, but what's interesting is that there are these emergent attentional concepts, these little mechanisms that have emerged to detect and control changes in attention, and they seem to just have come from the transformer attention head structure, which is interesting.
Now, if this is true, let's say that this multi-head attention structure creates these micro attention schemas, these little patterns of attention that are reusable and trained, then what does this mean? Well, first, we need to find out about the complexity, is it similar to human cognition? What happens if we put in hierarchies? Do we get more complex behavior?
And what's the actual relevance to social cognition?
So, I'm going to leave that for other people to think about it. I think it boils down partly to its engineering challenges, but I hope all of you consider both enjoying Keenan's talk about people who are doing this work in attention research, and also, when you think about your next projects, consider this direction.
Cool. Okay, we'll leave this. The main idea is that brains have different levels of organization, and the theories might will be the same thing.
Um definitely leave this.
Cool. Okay. Oh, yes, right. So, final final takeaways, thank you for bearing with me, Hikari.
So, how to build a theory of consciousness? You've got your big theory of everything. We were both very lucky to review papers for this for this conference, very privileged. Thank you for your submissions. And my takeaways to all of you are one, be clear about your material. What is the theory you're coming from? What are the objections? What are your angles?
Two, mechanism. What are you testing actually? What have you formalized? And please, what are the boundaries? Where does it start? Where does it end?
And final one, mystery.
What does this actually mean for consciousness if you have found something? You must link to behavioral data in humans, in animals, the kind of things that we feel intuitively might be important for consciousness.
Um I know that's vague, but I think that it's the thing that people say to you when you present a theory and they say, "Oh, does it do X?"
You should find a way to link your work to X. And that is how to build a theory that will not fall apart.
Thank you.
>> Amazing. Thank you Angie and Saba. Maybe we can take um one question. We went a little over.
>> Yes.
>> Thank you.
>> Thank you. Great talk.
>> Um so, naturally, I have two questions.
So, one is um but so, attention schema theory. Um this um so, this is actually about the bodily well, embodiment in a in a certain form, but um you mentioned Annelise's um essay, which comes from a different angle of bodily um bodily perception. So, I got a little bit confused over here if you are going to connect with Annelise's work and Annelise's path in attention schema theory.
>> Yeah, so I think the fundamental contradiction with Annelise's work and with biological naturalism is about the properties of information processing in biological systems. So, systems that are adaptive, autocatalytic, that are stochastic, right? That are multi-scale organization. All these things that we see in biology that we haven't currently necessarily built into frontier AI systems, but there are lots of people working on them in other ways.
And I think Annelise's argument is basically about emergence, that the information that we use when we think about qualia, it's not collapsible. It's all the way down, maybe down to the architecture, maybe down to the molecular level, maybe down to the quantum level.
Um so, the idea is that you can't separate the information that you have in your mind from the hard the wetware, the software that you're implementing cannot be separated from your physical body.
Does that make sense? Because your body is so complex and the information is interacting at so many scales.
But he hasn't proven this mathematically. He hired someone to work on it. I know Fernando.
Um they've sent out a theory, but no proof that it's true, biological naturalism is true. So, I await the answer.
>> Yeah, thank you for clarifying that. And then also the second question would be more about uh the the scope of the consciousness.
You just mentioned that I do really agree strongly that consciousness is also an umbrella term and we have to specific specify which type of consciousness we're talking about. Um so, in that sense, uh I just wanted to, you know, uh clarify {slash} confirm that your um you're trying to uh specify the accessible and functional parts of consciousness uh to the square of one more consciousness or are you like trying to uh stop the bridge between two separate two consciousnesses and and trying to deal the deal with those two parts, portions separate way.
>> Hm.
Well, we don't have any philosophers on team.
We would love to have more philosophers on our team. It is a central problem.
Yeah, please join us. You can email me.
What's my email? Um it's [email protected].
Um as for are we trying, I think our our idea was to take the premises of these theories in neuroscience and then the conclusions that they come to and then say, "Okay, if we test these mechanisms, do the same premises lead to the same conclusions?" And then we would extrapolate outwards to talk about phenomenology and behavior.
I did have a slide here about the like trying to interpret the phenomenology and the main thing that I found was that there's no consensus in philosophy about like what phenomenology is, what valence really means, how to find it, where it's located. Even biological accounts of valence actually like account for how the brain works, the things that we already know how it works, that is.
Don't seem to either agree or really fit the neuroscientific evidence.
So, this is sort of a problem.
And my main conclusion was that some of the philosophical definitions of qualia literally don't fit a human brain.
So, what do I do with that?
And I think it might be a case where we have to be a little bit post hoc.
Or we need people like France at the AMC and like Hikari who are willing to wade through the metaphysical challenges.
Yeah.
Thank you.
>> Thanks. Thank you.
>> [applause]
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23