Great read - looking forward to the rest of the series!
This reminded me of this recent-ish essay: https://gershmanlab.com/pubs/NeuroAI_critique.pdf. The core argument here (as I understand it) is that the very notion of biological inspiration for AI is somewhat ill-defined, because characterizing biological principles in the first place requires coming to the problem with a computational framework to make sense of the data: it’s not like the principles are just “there”. The proposal here is that rather than attempting to use neural/cognitive plausibility as a source of design principles, we can instead use it as a tiebreaker between candidate algorithms, under the assumption that algorithms that are actually implemented by biological systems are generally better.
As a separate but related note: even if we did have detailed, integrative, computational theories of complex cognitive phenomena, it’s not obvious that we should strive to integrate those insights into AI systems, because the constraints imposed on biological intelligence and AI systems are often quite different. For example, if a particular cognitive phenomenon is a byproduct of the fact that we humans have limited working memory, it’s not clear why we would try to bake that into AI, which is not bound by those same limitations. For some domains, the core problem being solved by both humans and AI may be similar enough that the solutions may be transferrable (and that doing so could be beneficial), but for others this seems less obvious.
Great points & pointers, thanks! On the different constraints, I certainly agree — but I tend to think that we'd like an ideal AI systems to be able to solve most problems at least as well as humans (though it might certainly outperform along some axes due to lacking some human constraints). Thus, in the cases where humans still appear to be more efficient or better continual learners etc., it seems plausible that there is something to be something to be learned at the computational level at least.
I hope you will address what seems to me the core question about the relationship between AI and cognitive science, which is “Does the brain do anything resembling backprop and gradient descent” when humans learn by exposure to pattern-rich sensory inputs.
I “liked” your reply because anyone with a quick answer to this query would be making stuff up. I am not a cognitive scientist, but I can read the cog sci and neuroscience literature when I have to, and my impression is that it is all over the map on this central question.
I think an important problem with the field of NeuroAI is that strong claims are too often made based on extremely weak evidence. For instance, you write: " language models can predict human imaging data remarkably well", but it is important to be clear what the authors you cite found. They reported models could account for near 100% of explainable variance, but the explainable variance was very very low (in most cases between 4-10%), and that a similar amount of variance was observed in non-language areas, raising questions about how to interpret the findings. And you write cognitive scientists have suggested that AI progress refutes one of the dominant linguistic paradigms. Again, that is true, but the evidence for this claim is extremely weak, as reviewed here. https://psycnet.apa.org/record/2026-83323-001. Do you think the papers you cite make a strong cases for ANN-human alignment, and challenge the role of innate priors? I think the field could use a bit more scepticism.
I generally agree that some skepticism is warranted — as you probably know, I've written about computational challenges to trusting some model-brain representational comparisons myself, (e.g. https://arxiv.org/abs/2507.22216). That theme will maybe come up more strongly in future posts (including the one about what cog sci can learn from AI). But I also think that the fact that e.g. explainable variance is low is in part due to the critique I make here: explainable variance is low because neuroscience tasks are just not varied, complex, or sequential enough to drive neural activity all that strongly (or allow us to account for individual differences all that well).
On your specific points for skepticism, I think some are valid. But I also think you read the human-side literature far too uncritically. The "impossible" languages results are based on a couple of almost-anecdotal studies done on *adults* who were already fluent in another language. They would have to be done in the critical period and as first language to be a really meaningful measure of what's learnable. I'd love to see someone do a modern large-scale developmental study on that, and a comparative study on models, which do show some language inductive biases as you probably know: https://arxiv.org/abs/2401.06416
I think the impossible language article you cite is a good case in point. The simulations the authors report do not support the claims the authors make. Here is a commentary I submitted for BBS paper that reviews the findings on impossible languages and LLMs. https://arxiv.org/pdf/2511.11389. Note, one of the impossible language simulations we published in our original work would not be learnable as it requires a super-human STM: https://aclanthology.org/2020.coling-main.451/. And as noted in our commentary on the LLM Centaur that was claimed to account for 159 out of 160 experiments better than cognitive models, including STM model, Centaur could sometimes repeat 256 digits perfectly: https://doi.org/10.31234/osf.io/v9w37_v4. All examples of what I'm talking about. Sure, there are problems with the cognitive science literature as well, but the claims in NeuroAI are being uncritically published in all the top journals. Criticisms of these claims, not so much.
I don't think you're really engaging with my point about the definition of impossible languages being based on tiny experiments on adults in short time, except possibly insofar as you mention something requires super-human STM — but even that might be different if the humans were experts at the language from learning from infancy (just as chess experts' STM capacity for board positions shouldn't be naively extrapolated from novices).
For your broader point, I think it's easy to pick cases where a model fails and then claim it's bad. I have plenty of my own qualms with Centaur. But the fact it is a much better model than any other to date of how behavior varies across tasks, and that is an important feature in itself. No model is perfect.
And where there is some failure of a model to capture some important human behavior, and you identify and show how to fix that failure (without making the model worse at everything else), you can definitely get into top journals (e.g., https://www.nature.com/articles/s41586-025-09631-6 — where I was surprised at the tier of journal it got accepted in). I personally find work much more compelling when it doesn't just say "this is bad," but points in a direction which is better.
With regards to impossible languages, it is an interesting question as to whether unattested languages that violate UG are unattested because they are hard to learn by humans (as argued by Chomsky) or for some other unexplained reason. But it is important to note that Kallini et al. challenged innate language priors based on simulations showing that LLMs *do* find impossible languages more difficult to learn. By contrast, Mitchell and Bowers 2020 showed that LLMs find impossible languages easy to learn, and in the BBS comment article that I linked to (https://arxiv.org/pdf/2511.11389), we explain why Kallini et al. drew the wrong conclusion. Of course, the argument can be changed to claim that LLMs are like humans: maybe both humans and LLMs learn impossible languages equally easily. But that was not the claim of anyone thus far, there is no human empirical evidence for this, and the current findings most certainly to not challenge Chomsky.
Regarding Centaur, the model has no human-like mechanism for generating RTs (and accordingly can output RTs of 1ms while maintaining accuracy), and it has a STM approximately 1 order of magnitude larger than humans (and can apply this capacity to any task it confronts). And it provides not explanation for any cognitive phenomena (why do humans have a STM of about 4 items, why does it take about 500 ms to respond to a stimulus?, etc.). I quite like Shanahan’s characterisation: that LLMs are good role playing (except when you test them off script). As such, they tell you little about the mechanisms of the mind, the point of models in cognitive science.
Of course, it is great to build models. But building models and making causal claims based on correlational studies where you predict unexplained variance is a strategy that will lead to many unwarranted claims. The fact that it is hard to publish studies that show that the published claims are not justified after running experiments that test hypotheses is a serious problem with the NeuroAI field. A series of articles (not only ours) were all rejected from Nature that all showed that Centaur is not at all human-like, and despite its ability to account for unexplained variance in 160 experiments, not an improvement on cognitive models that precede it. Indeed, it is worse, as it does not even attempt to explain anything (other than the claim that human cognition is based on predictions?).
I will end with two basic points:
First, falsification is important. It is enough to (compellingly) show that important claims are false or unwarranted without providing an alternative model. Falsifications should not be relegated to some specialised psychology journal while the unwarranted claims are published in Nature and Science and the like. Indeed, most researchers in NeuroAI largely ignore the psychology literature.
Second, NeuroAI has to start adopting standard methods of science and stop with the prediction of unexplained variance on benchmarks.
Well the other problem is that after the industrial revolution and information revolution we have been fixated on computational architectures and assumed that the brain must follow a similar structure, same happened with quantum computers then engineers have been searching for quantum effects in brains. It's a well known cognitive fallacy that we don't seem to be able to control...
I do think that our need to analogize between whatever complex system we have on hand and the brain is both a weakness and an interesting fact in itself about human cognition. That doesn't mean all the analogies are bad, though.
For the record though, I'm pretty skeptical of any quantum effects in the brain being relevant at a cognitive scale, even though the difference between quantum computers and non-quantum computers is one of the interesting cases where there is really an effect of an implementation-level feature all the way up to fundamental computational questions. The amount of noise in a quantum system just seems way too high for reliability at scale at even a few degrees K, let alone body temperature.
Yes consider that it all started with phrenology we didn't do so bad and as you pointed out analogies could be the manifestation of some form of transcendental laws maybe operating at different scales or substrates.
I was reading about an interesting theory suggesting the brain could be an analog interference engine... I like those odds!
On quantum I am also very skeptical for the same reasons you mentioned.
Great read! I have always wondered why there is a rush to create systems that are gargantuan and consume lots of resources when we are tiny and consume so little. You provide lots of good reasons for that. Nevertheless, I am an HI (human intelligence) optimist. AI is great but what we will do with these new tools is even greater.
The point about fragmentary understanding breaking down at integration points feels key — many AI failures seem to emerge not from missing components, but from how multiple constraints interact.
This is a great argument to present to folks who state so confidently that the current approaches are wrong because that is totally not what our mind is doing.
I want to help. As the least likely person to contribute to the field of AI in the early 2022, I read a lot of books and papers, watched multiple videos in philosophy, cogsci, AI, psychology, etc. I learned about basic observations the "big picture" needs to explain.
I had many useful early insights, but only recently I could connect many dots to believe that I can help. And of all the AI labs, I think GoogleDeepMind is the best one to understand. You had MCTS. I propose Semantic Binary Search. You can appreciate its potential. But you are right - a big demo is needed and I am only a one-man team.
There are many misleading assumptions to throw away. Fine lines between some of them and the better ones. But it's worth it. The promise is there!
You are welcome to my Substack. Also, consider taking a look at this paper where I take the best points so far and connect many dots - https://philpapers.org/rec/NAUNOC-2
> And more generally, I’d argue that the current dominant paradigm in AI has not primarily grown out of approaches that designs systems on the basis of cognitive science.
So true. Yet some make great progress.
> building in how we think we think does not work in the long run.
Yes for practical AI, but not for theoretical AI. e.g
> Thus, I think that the fundamental reason that cognitive science has not contributed more to engineering modern AI systems is that our current understanding of cognition is often too fragmentary or abstract.
The article poses an important question, but I am not sure it is well defined.
Are there any principles in Cognitive Science?
We don't know the principles of natural intelligence (yet) and even if we did, they may not even help us to build machines that think.
The various branches of neuroscience have produced an abundance of theories, but we don't know how to identify the ones that are correct or the ones that are incorrect. To prove my point: Since Lord Adrian's Nobel work, thousands of theories have been published on how the visual system works (for different species), and still, you can happily publish a new one every year.
Is there sich a thing as "modern" AI or are we talking about a steady stream of insights since the 1940s?
Finally, is there any technology at all that is built around the principles of an animal?
Planes certainly don't flap their wings and ships don't use their fins for propulsion. Even robots don't use muscle-tendon actuators.
Just my two cents as a computational neuroscientist.
The answer is very simple. The A.I. field engages mostly in what can be called engineering - it is, for all intents and purposes, extremely bottom-up in how it operates, defines, and refines itself. It's not conducting what would be known as science - which includes such non-trivial things as evidence synthesis, theorizing, hypothesis testing and falsification etc. Task and experiment creation, for example, have proven to be as you note extremely challenging for the field. This is not random. Task and experiment creation are extremely challenging for cognitive science and behavioral science, as well. It is in this process of arduous science that things get developed - they don't "emerge" from compute and number of people thrown at the problem, they get developed as a result of iterative hypothesis testing and painstaking top-down theorizing that aims to be mechanistic in nature, specification of alternative hypotheses etc. The A.I. field pretends it did not originate in the 1950s at the merger of behavioral sciences, psychology, linguistics and everyone else vaguely interested in information processing and systems-level frameworks for studying it. It pretends it originated as a God-like extension of some genius polymath ML offshoot. There are consequences for not remembering your roots. And engaging with it substantively requires non-GPT-mediated understanding of very complex literature. So they just skipped it ;)
This is.. alchemy pretending to be chemistry, in essence. By choice ;)
Great read - looking forward to the rest of the series!
This reminded me of this recent-ish essay: https://gershmanlab.com/pubs/NeuroAI_critique.pdf. The core argument here (as I understand it) is that the very notion of biological inspiration for AI is somewhat ill-defined, because characterizing biological principles in the first place requires coming to the problem with a computational framework to make sense of the data: it’s not like the principles are just “there”. The proposal here is that rather than attempting to use neural/cognitive plausibility as a source of design principles, we can instead use it as a tiebreaker between candidate algorithms, under the assumption that algorithms that are actually implemented by biological systems are generally better.
As a separate but related note: even if we did have detailed, integrative, computational theories of complex cognitive phenomena, it’s not obvious that we should strive to integrate those insights into AI systems, because the constraints imposed on biological intelligence and AI systems are often quite different. For example, if a particular cognitive phenomenon is a byproduct of the fact that we humans have limited working memory, it’s not clear why we would try to bake that into AI, which is not bound by those same limitations. For some domains, the core problem being solved by both humans and AI may be similar enough that the solutions may be transferrable (and that doing so could be beneficial), but for others this seems less obvious.
Great points & pointers, thanks! On the different constraints, I certainly agree — but I tend to think that we'd like an ideal AI systems to be able to solve most problems at least as well as humans (though it might certainly outperform along some axes due to lacking some human constraints). Thus, in the cases where humans still appear to be more efficient or better continual learners etc., it seems plausible that there is something to be something to be learned at the computational level at least.
I hope you will address what seems to me the core question about the relationship between AI and cognitive science, which is “Does the brain do anything resembling backprop and gradient descent” when humans learn by exposure to pattern-rich sensory inputs.
Not in the next few, but at some point it definitely deserves a post — once I've caught up with the literature :)
I “liked” your reply because anyone with a quick answer to this query would be making stuff up. I am not a cognitive scientist, but I can read the cog sci and neuroscience literature when I have to, and my impression is that it is all over the map on this central question.
I think an important problem with the field of NeuroAI is that strong claims are too often made based on extremely weak evidence. For instance, you write: " language models can predict human imaging data remarkably well", but it is important to be clear what the authors you cite found. They reported models could account for near 100% of explainable variance, but the explainable variance was very very low (in most cases between 4-10%), and that a similar amount of variance was observed in non-language areas, raising questions about how to interpret the findings. And you write cognitive scientists have suggested that AI progress refutes one of the dominant linguistic paradigms. Again, that is true, but the evidence for this claim is extremely weak, as reviewed here. https://psycnet.apa.org/record/2026-83323-001. Do you think the papers you cite make a strong cases for ANN-human alignment, and challenge the role of innate priors? I think the field could use a bit more scepticism.
I generally agree that some skepticism is warranted — as you probably know, I've written about computational challenges to trusting some model-brain representational comparisons myself, (e.g. https://arxiv.org/abs/2507.22216). That theme will maybe come up more strongly in future posts (including the one about what cog sci can learn from AI). But I also think that the fact that e.g. explainable variance is low is in part due to the critique I make here: explainable variance is low because neuroscience tasks are just not varied, complex, or sequential enough to drive neural activity all that strongly (or allow us to account for individual differences all that well).
On your specific points for skepticism, I think some are valid. But I also think you read the human-side literature far too uncritically. The "impossible" languages results are based on a couple of almost-anecdotal studies done on *adults* who were already fluent in another language. They would have to be done in the critical period and as first language to be a really meaningful measure of what's learnable. I'd love to see someone do a modern large-scale developmental study on that, and a comparative study on models, which do show some language inductive biases as you probably know: https://arxiv.org/abs/2401.06416
I think the impossible language article you cite is a good case in point. The simulations the authors report do not support the claims the authors make. Here is a commentary I submitted for BBS paper that reviews the findings on impossible languages and LLMs. https://arxiv.org/pdf/2511.11389. Note, one of the impossible language simulations we published in our original work would not be learnable as it requires a super-human STM: https://aclanthology.org/2020.coling-main.451/. And as noted in our commentary on the LLM Centaur that was claimed to account for 159 out of 160 experiments better than cognitive models, including STM model, Centaur could sometimes repeat 256 digits perfectly: https://doi.org/10.31234/osf.io/v9w37_v4. All examples of what I'm talking about. Sure, there are problems with the cognitive science literature as well, but the claims in NeuroAI are being uncritically published in all the top journals. Criticisms of these claims, not so much.
I don't think you're really engaging with my point about the definition of impossible languages being based on tiny experiments on adults in short time, except possibly insofar as you mention something requires super-human STM — but even that might be different if the humans were experts at the language from learning from infancy (just as chess experts' STM capacity for board positions shouldn't be naively extrapolated from novices).
For your broader point, I think it's easy to pick cases where a model fails and then claim it's bad. I have plenty of my own qualms with Centaur. But the fact it is a much better model than any other to date of how behavior varies across tasks, and that is an important feature in itself. No model is perfect.
And where there is some failure of a model to capture some important human behavior, and you identify and show how to fix that failure (without making the model worse at everything else), you can definitely get into top journals (e.g., https://www.nature.com/articles/s41586-025-09631-6 — where I was surprised at the tier of journal it got accepted in). I personally find work much more compelling when it doesn't just say "this is bad," but points in a direction which is better.
With regards to impossible languages, it is an interesting question as to whether unattested languages that violate UG are unattested because they are hard to learn by humans (as argued by Chomsky) or for some other unexplained reason. But it is important to note that Kallini et al. challenged innate language priors based on simulations showing that LLMs *do* find impossible languages more difficult to learn. By contrast, Mitchell and Bowers 2020 showed that LLMs find impossible languages easy to learn, and in the BBS comment article that I linked to (https://arxiv.org/pdf/2511.11389), we explain why Kallini et al. drew the wrong conclusion. Of course, the argument can be changed to claim that LLMs are like humans: maybe both humans and LLMs learn impossible languages equally easily. But that was not the claim of anyone thus far, there is no human empirical evidence for this, and the current findings most certainly to not challenge Chomsky.
Regarding Centaur, the model has no human-like mechanism for generating RTs (and accordingly can output RTs of 1ms while maintaining accuracy), and it has a STM approximately 1 order of magnitude larger than humans (and can apply this capacity to any task it confronts). And it provides not explanation for any cognitive phenomena (why do humans have a STM of about 4 items, why does it take about 500 ms to respond to a stimulus?, etc.). I quite like Shanahan’s characterisation: that LLMs are good role playing (except when you test them off script). As such, they tell you little about the mechanisms of the mind, the point of models in cognitive science.
Of course, it is great to build models. But building models and making causal claims based on correlational studies where you predict unexplained variance is a strategy that will lead to many unwarranted claims. The fact that it is hard to publish studies that show that the published claims are not justified after running experiments that test hypotheses is a serious problem with the NeuroAI field. A series of articles (not only ours) were all rejected from Nature that all showed that Centaur is not at all human-like, and despite its ability to account for unexplained variance in 160 experiments, not an improvement on cognitive models that precede it. Indeed, it is worse, as it does not even attempt to explain anything (other than the claim that human cognition is based on predictions?).
I will end with two basic points:
First, falsification is important. It is enough to (compellingly) show that important claims are false or unwarranted without providing an alternative model. Falsifications should not be relegated to some specialised psychology journal while the unwarranted claims are published in Nature and Science and the like. Indeed, most researchers in NeuroAI largely ignore the psychology literature.
Second, NeuroAI has to start adopting standard methods of science and stop with the prediction of unexplained variance on benchmarks.
Well the other problem is that after the industrial revolution and information revolution we have been fixated on computational architectures and assumed that the brain must follow a similar structure, same happened with quantum computers then engineers have been searching for quantum effects in brains. It's a well known cognitive fallacy that we don't seem to be able to control...
I do think that our need to analogize between whatever complex system we have on hand and the brain is both a weakness and an interesting fact in itself about human cognition. That doesn't mean all the analogies are bad, though.
For the record though, I'm pretty skeptical of any quantum effects in the brain being relevant at a cognitive scale, even though the difference between quantum computers and non-quantum computers is one of the interesting cases where there is really an effect of an implementation-level feature all the way up to fundamental computational questions. The amount of noise in a quantum system just seems way too high for reliability at scale at even a few degrees K, let alone body temperature.
Yes consider that it all started with phrenology we didn't do so bad and as you pointed out analogies could be the manifestation of some form of transcendental laws maybe operating at different scales or substrates.
I was reading about an interesting theory suggesting the brain could be an analog interference engine... I like those odds!
On quantum I am also very skeptical for the same reasons you mentioned.
Great read! I have always wondered why there is a rush to create systems that are gargantuan and consume lots of resources when we are tiny and consume so little. You provide lots of good reasons for that. Nevertheless, I am an HI (human intelligence) optimist. AI is great but what we will do with these new tools is even greater.
This framing resonates.
The point about fragmentary understanding breaking down at integration points feels key — many AI failures seem to emerge not from missing components, but from how multiple constraints interact.
This is a great argument to present to folks who state so confidently that the current approaches are wrong because that is totally not what our mind is doing.
wrote abt this along the same lines a couple of years ago, which might be of interest: https://www.aishwaryadoingthings.com/from-physics-envy-to-biology-envy
Thanks for this great outline!
I want to help. As the least likely person to contribute to the field of AI in the early 2022, I read a lot of books and papers, watched multiple videos in philosophy, cogsci, AI, psychology, etc. I learned about basic observations the "big picture" needs to explain.
I had many useful early insights, but only recently I could connect many dots to believe that I can help. And of all the AI labs, I think GoogleDeepMind is the best one to understand. You had MCTS. I propose Semantic Binary Search. You can appreciate its potential. But you are right - a big demo is needed and I am only a one-man team.
There are many misleading assumptions to throw away. Fine lines between some of them and the better ones. But it's worth it. The promise is there!
You are welcome to my Substack. Also, consider taking a look at this paper where I take the best points so far and connect many dots - https://philpapers.org/rec/NAUNOC-2
I am willing to help!
Sounds right. I'm looking forward to reading your follow-up posts.
Cog sci is a deeply dated, reduced framework. TIme for neurobiology to take the lead, and to put AI in the dustbin.
> Historically, cognitive science and AI were tightly coupled fields,
In my reading cognitive psychologists pay lip service to AI but mostly read their empirical journals, often very specialized.
Cognitive science is supposed to be the interdisciplinary study of the mind. But that again is lip service.
I outline an approach here that combines architecture-based AI with broad and deep cognitive science: [A Manifesto for an Integrative Design-oriented Approach to Understanding Humans as Autonomous Agents – CogZest](https://cogzest.com/projects/a-manifesto-for-integrative-design-oriented-cognitive-science-and-ai/)
> We know a lot about the brain across many levels of analysis.
yes and no. mostly no.
> arguing that AI is missing core ingredients of human intelligence
And I answered here [Why You Can't Say AI Is—or Is Not—Intelligent](https://luccogzest.substack.com/p/why-you-cant-say-ai-isor-is-notintelligent).
> And more generally, I’d argue that the current dominant paradigm in AI has not primarily grown out of approaches that designs systems on the basis of cognitive science.
So true. Yet some make great progress.
> building in how we think we think does not work in the long run.
Yes for practical AI, but not for theoretical AI. e.g
1. [A Manifesto for an Integrative Design-oriented Approach to Understanding Humans as Autonomous Agents – CogZest](https://cogzest.com/projects/a-manifesto-for-integrative-design-oriented-cognitive-science-and-ai/) and
2. [Aaron Sloman 1993 Prospects for AI as the General Science of Intelligence Aaron Sloman](https://cogaffarchive.org/Aaron.Sloman_prospects.pdf).
3. And [My Ph.D. thesis: Goal Processing in Autonomous Agents](http://www.cs.bham.ac.uk/research/projects/cogaff/Luc.Beaudoin_thesis.pdf): basically doing theoretical AI.
> However, in my experience cognitive (neuro)science has focused most on studying intelligence within small sets of relatively simplified tasks.
yes!
> build tasks that test only the capability we’re interested in, in settings where that capability is the only perfect solution.
again contrast [A Manifesto for an Integrative Design-oriented Approach to Understanding Humans as Autonomous Agents – CogZest](https://cogzest.com/projects/a-manifesto-for-integrative-design-oriented-cognitive-science-and-ai/).
> Thus, I think that the fundamental reason that cognitive science has not contributed more to engineering modern AI systems is that our current understanding of cognition is often too fragmentary or abstract.
I can't find much to disagree with you on.
The article poses an important question, but I am not sure it is well defined.
Are there any principles in Cognitive Science?
We don't know the principles of natural intelligence (yet) and even if we did, they may not even help us to build machines that think.
The various branches of neuroscience have produced an abundance of theories, but we don't know how to identify the ones that are correct or the ones that are incorrect. To prove my point: Since Lord Adrian's Nobel work, thousands of theories have been published on how the visual system works (for different species), and still, you can happily publish a new one every year.
Is there sich a thing as "modern" AI or are we talking about a steady stream of insights since the 1940s?
Finally, is there any technology at all that is built around the principles of an animal?
Planes certainly don't flap their wings and ships don't use their fins for propulsion. Even robots don't use muscle-tendon actuators.
Just my two cents as a computational neuroscientist.
The answer is very simple. The A.I. field engages mostly in what can be called engineering - it is, for all intents and purposes, extremely bottom-up in how it operates, defines, and refines itself. It's not conducting what would be known as science - which includes such non-trivial things as evidence synthesis, theorizing, hypothesis testing and falsification etc. Task and experiment creation, for example, have proven to be as you note extremely challenging for the field. This is not random. Task and experiment creation are extremely challenging for cognitive science and behavioral science, as well. It is in this process of arduous science that things get developed - they don't "emerge" from compute and number of people thrown at the problem, they get developed as a result of iterative hypothesis testing and painstaking top-down theorizing that aims to be mechanistic in nature, specification of alternative hypotheses etc. The A.I. field pretends it did not originate in the 1950s at the merger of behavioral sciences, psychology, linguistics and everyone else vaguely interested in information processing and systems-level frameworks for studying it. It pretends it originated as a God-like extension of some genius polymath ML offshoot. There are consequences for not remembering your roots. And engaging with it substantively requires non-GPT-mediated understanding of very complex literature. So they just skipped it ;)
This is.. alchemy pretending to be chemistry, in essence. By choice ;)
Great question! This (relatively accessible, I think) paper is a good place to start for an academic treatment: https://onlinelibrary.wiley.com/doi/full/10.1111/j.1756-8765.2010.01116.x
Jay (the author of that) also just co-wrote a less-academic book on the subject:
https://www.amazon.com/Emergent-Mind-Intelligence-Arises-Machines/dp/1541605268/