Mainstream Criticism on Education
Isn't it rather amusing how uncontroversial and common claiming that "education is wrong and inefficient" has become? Still... isn't it a big concern when people scream from the top of their lungs about a certain problem, yet nothing gets fixed? After all, it's absurd and pessimistic to think that societal problems aren't solvable, it might be just that we don't have the right knowledge to solve them yet... But how can it be the case, when it feels as if countless resources have gone into changing education and yet we still complain about it. Maybe, like in many cases, the crowd is simply wrong... What if education is good enough the way it is? Could it be that children are simply too lazy? Learning is, as they say, supposed to be hard! Or is it?
It's hard to tell whether these questions really yield any clear responses. Which question should we follow first? Which one is epistemically more useful? You'd be circling around what the current system does and play within its own framework.
I think, due to how old these institutions are, and how much they have infiltrated our culture simply through existing for centuries (where we now see schooling as a given right, as something that has always existed), the assumptions presented by the educational system about what learning is have become so deeply ingrained in our brains that we take most of them for granted. Answering most of the above questions would therefore entail bringing in assumptions from the very educational system we are trying to examine. I therefore think the best way to understand the state of things is by starting from first principles and, to the extent we can, thinking about what Education Is In Itself!
What is the objective of schooling?
Let us assume that the core reason schools exist is to educate pupils [1] (although you can argue that’s not their primary function). Education in its broader definition is the deliberate attempt to help the subject develop knowledge.
<Throughout this essay I’ll be referring to knowledge in the broader definition that David Deutsch offers as information with causal properties>
To make it even simpler, education is the environment which facilitates the acquisition of knowledge. That encompasses all tools, methods and external agents (what schooling attempts to be if we take its core purpose that of educating).
The process through which the acquisition of knowledge happens at an individual level is called learning. It’s not an external process to be controlled, but rather an internal one where the subject’s understanding of the world shifts and adapts so that it can better predict the future. Education is the environment which facilitates the learning process, and although it cannot control how each individual acquires knowledge, it can create the premises that speed up and strengthen that process. It therefore follows that Education (the environment which includes all the tools, methods, agents; teachers, classrooms, curriculum) is downstream of learning, and a way to improve it is by building upon the principles of whatever learning entails.
Learning Is Model Change
What actually switches when you learn?
Learning is a process through which what you, as a human being, are capable of changes. We can learn by mimicking what others do, or retain information without developing any deeper understanding of it, but throughout this essay, when I refer to learning, I mean something slightly narrower: the acquisition of generative knowledge, where knowledge is information with causal properties. It's what allows us to shape our environment, not only be shaped by it. In other words, I am interested in learning that changes not merely what you can recall or reproduce, but what you are capable of explaining, predicting and creating. Throughout this essay, I refer to these capabilities of explanation, prediction and creation as the generative functions of your mental model.
Using this framework, dreams serve as a useful illustration of the brain's generative character: they are creations of the generative model using learned weights plus encoded context, or versions thereof. The core difference between experiencing a dream and experiencing the world while awake is that waking cognition is much more tightly constrained by sensory input. This lack of grounding in dreams can produce inconsistencies with reality; hallucinations, if you will. The idea I want to convey is that our mental model is responsible for how we interpret reality. While we are all grounded in roughly the same inputs and constraints, we all have different generative models.
Our model generates expectations about what will happen, and those expectations can be wrong. It is this capacity (to anticipate the short-term future, predict the consequences of our actions, and to some extent, over the past few centuries, the consequences of those consequences) that helped us survive. We must take into account, however, that this capability of predicting the future using your predictive model, despite being extremely useful, doesn't mean it's true (because it doesn't really have to stay true, learning is a means to an end not an end in itself, and that end is solving problems). It helps us only to the extent of the problems it can solve.
The reason why inertia might be quite a counterintuitive concept for kids is that their world model finds it hard to explain it. Frankly speaking, it doesn't even need a reason for its existence, since in most kids' day-to-day experiences, objects around them (like toy cars) don't just continue moving at infinitum. Therefore, the lack of such a concept is coherent with their existing world model (such a concept could probably threaten some kids' explanations about the world they currently hold) in which things slow down unless you push on them and where everything naturally falls down. Unless they encounter a moment where their explanations don't hold, their world model serves them well enough and has no reason to change. Such a moment that makes them question their model could be skating on ice, since you can move quite a long distance once you get a single push, without stopping for a long time, unless you got some form of obstacle in your way or you deliberately choose to slow down using certain maneuvers…
Here's where kids' curiosity comes into place. Whenever they encounter something their model couldn't explain, suddenly, a question becomes possible: why can I slide so far on ice but barely at all on grass? And then, some of the time, they can slide on grass when it's dewy in the morning.
Here is where curiosity and knowledge acquisition enter the picture. Knowledge becomes necessary because there is now a problem. Something happened that the learner's existing model cannot satisfactorily explain. We can call the difference between what the model expected and what reality permitted a prediction gap. But a failed prediction is only one species of a more general thing: a problem.
Problems
Not one of those things you do on a test or at the end of a lesson.
A problem can be a contradiction between explanations, an unexplained phenomenon, an unsatisfied goal, or a situation in which the learner suddenly realizes that their existing model cannot do something they need it to do—a prediction gap of some sort.
Problems have been advertised in schools as mere features of a lesson, mostly to help the teacher assess his students, when in reality problems are the origin of meaningful learning. No such learning could happen without the existence of a prediction gap, since knowledge's purpose is to fill in these gaps and improve our mental models. The right problem makes the prediction gap visible to the student; it gives the lacking knowledge a shape and turns an unknown unknown into a known unknown.
The moment your mental model of the world doesn't hold up in front of some situation, something in your assumptions or reasoning is wrong, and as a result you must come up with a better guess for how to predict things.
The good thing is you never have to start with the truth. Problems don't deal or care about truth; they only care about better solutions. Therefore, you should be merely satisfied with trying to reach for it more every single time. Throughout history, progress hasn't been made when people were looking for the absolutely correct answer, but rather the better one. We usually started with crude local explanations that were predicting our environment well enough. Whenever our environment changed, or we started doing things differently, the frameworks driving our prediction models collapsed because they weren't serving us anymore, so we had to develop better explanations. Our explanations tend to start from crude observations, but as we explore and experience more of the world, we encounter problems which need new explanations to create a more coherent tapestry of how the world works.
Back to kids' models of the world: objects ceasing their motion when a continuous force isn't applied, things falling down, or Santa Claus delivering presents on Christmas Eve are all valid predictions that satisfy the local model they hold. As long as they don't encounter new problems, there won't be a need for a new mental model. The geocentric model explained to people why the Sun and Moon were continuously in motion relative to us; it served them well enough until we needed a better model to reach for new solutions to new problems that revealed gaps in our ways of thinking. The need for new mental models arises when ours cannot explain certain phenomena, which happens as we tend to expand more globally and search for more all-encompassing and accurate explanations due to the increasing difficulty of our problems.
By "globally more accurate" explanations, I mean types of explanations that can withstand way more criticism and are more universal than their predecessors. These types of explanations are not just helpful in the sense that they help us solve new problems; they also help us discover new ones.
Take Newtonian physics for example: was it explaining reality as it unfolded? No! No theory explains reality 1:1, though some explain it better than others. But all in all, Newton's concepts helped us make predictions enough so we could shape the environment around us and consider them way more useful than any others we had at the time. They helped us solve countless problems we didn't conceive of solving back then and also expanded not just the scope of new problems we didn't know existed till then, but also our ways of thinking about problems.
This brings us to the idea that knowledge compounds; as it expands, it allows for a better understanding of our previous mental models, strengthens the connections between existing concepts, discovers gaps in previous modes of thought, allows us to solve new problems, adds new constraints, and as a result facilitates even more discovery of newer knowledge. Always though, knowledge's meaning comes from the existence of a problem. Einstein realized something was wrong with Newton's framework because Newtonian mechanics and Maxwell's electrodynamics could not both remain true under the same assumptions about space, time, and motion. It was the contradiction that created the problem which later birthed Relativity.
This is one of the central parallels of the essay. At the scale of civilization, a problem exposes the limitation of an existing explanation and forces us to construct a better one. At the scale of an individual learner, the novelty is different because civilization may already know the answer, but the learner's model faces the same structural problem: its current explanation cannot yet generate what the new explanation would allow it to generate. Scientific discovery and individual learning are not identical events, but they may share the same underlying logic of knowledge growth.
Guesses
Or how we solve any problem
How does one start wrestling with a problem? Where are you even supposed to start searching for a solution you don't know yet?
Any solution of a problem starts as a simple guess, also referred to as a “conjecture”. A conjecture is an active task of searching the latent space of ideas you have to possibly come up with a better explanation as to “What was wrong in your current model?” and from there trying to find a better explanation that would be more universal in the sense that it would explain what the previous model did plus what it couldn’t.
Let me go a bit deeper first and explain that, by a guess/conjecture, what we truly mean is a coherent story that you construct and then expose to criticism from your other conjectures, observations, experiments, and the consequences it predicts. By story I don’t mean fiction in the traditional sense, but rather a model of how you think something works. Imagine I kick a football and it shoots forward. I might tell myself that the kick has put some quantity of “motion” into the ball, which it gradually uses up as it rolls until eventually there is none left and the ball stops. That story is wrong, but it is not stupid. It is internally coherent, it seems to explain what I just saw, and I can use it to make predictions. If I kick another ball tomorrow, I expect a harder kick to give it more motion and therefore make it travel farther. The story can survive until I find a problem with it, perhaps because I encounter something it cannot explain.
Another example could be astrology, which often makes predictions vague enough to be compatible with almost any outcome, so they are difficult to actually prove false. This is partly why such stories can survive for so long: they can remain coherent with someone’s other models of the world while constantly being reinterpreted around whatever happens. The distinction is that a good conjecture which has the chance of becoming a plausible explanation has to make a claim that can be falsified, be internally coherent, and not be arbitrarily adjustable whenever reality disagrees with it. So while we don’t logically derive these stories from observations of the real world, they are supposed to tell us, to an approximation, how the real world behaves. The real world acts more or less as a constraint on all the possible stories we can tell ourselves because it allows us to rule out explanations that simply cannot survive contact with it. We can avoid doing this, obviously, by having stories that never really say anything about the real world, and so, through their vagueness, they survive precisely because they aren’t falsifiable in the first place.
A story can be perfectly coherent inside your head, but the moment it predicts something that nature refuses to do, it loses its potential of becoming a good explanation, or if it predicts nothing at all, it simply remains what it is, a story...
One thing to note is that certain problems which are harder to solve cannot conceive a solution with only one set of conjectures, since they might require additional explanations that we have yet to find, which refers back to our idea of “knowledge compounds”. Therefore it’s hard to believe one can come up with nuclear fusion without first knowing chemistry in the first place…
This brings us to the question of what problems are worth solving at any given time and how do you find plausible and testable conjectures that can amount to an explanation?
Starting with the first question, the answer is to look, usually, at the edge of the newest innovation (this applies both on a civilization scale as well as on an individual one), since that’s the least tested part of our current model (due to its infancy) and that’s where you’re most likely to find gaps and new problems (because if we had a complete model at that edge, why would we have an edge in the first place?).
Addressing the latter question, guessing can take a lot of time, an infinite amount, if done randomly. The question is how does one know where to search in the vast latent space? Is there a filter that allows us to speed up this process? That brings us to the next part of this essay.
Constraints
Dos and Don'ts
Since the latent space where you’re searching for guesses is infinite, the obvious method one can use to shrink it is by excluding the somewhat wrong parts. But what can even count as “wrong” when searching for a framework that can solve a problem we don’t know the answer to yet?
The core constraints come from the problem itself. By choosing one problem over another, you ignore an infinity of other areas of the latent space. The problem gives the space a rough shape based on your existing knowledge, which again acts as a constraint. As such, we come to comprehend the importance of understanding the problem not just as a way of acknowledging the gap between our predictive model and reality but also as a method of getting closer to a more accurate explanation.
Constraints can also come from external agents, textbooks, nature, previous explanations, sensorial inputs, memories, or other sources outside your own scope. In other words, constraints are anything that allows us to rule out regions of the latent space and point our search towards the areas where a useful conjecture might actually exist.
What this search for good explanations looks like is close to a form of triangulation in a high-dimensional space: we utilise our existing knowledge to estimate where new knowledge might lie, allowing knowledge to grow from within. Yet the reach of our knowledge extends beyond what it can already explain. A theory does not only tell us what we know; by colliding with problems, it also exposes questions whose answers we do not yet possess. What was previously an unknown unknown can thereby become a known unknown, giving conjecture a new region in which to search. A student can begin to see why derivatives matter when you give them a problem where the velocity of a car is variable and ask what is happening at one particular instant. The problem creates a hole, they realise their model was incomplete, and now they formulate a question; which turns out to be half of the way to understanding the solution. We should stop treating students as clueless containers into which information must simply be poured. One can simply let the knowledge they already possess expose what is missing, and then give them the constraints they need to search for it.
Our knowledge cannot always expose these unknown unknowns, primarily if it cannot yet frame the problem and understand it well enough, which introduces the need for external instructions, these serve as catalysts that help the subject reduce the search space faster.
As said previously, constraints' role is to limit our search for an explanation, they can be learned and implied by our existing knowledge or be externally supplied. We will take a short detour into the world of Artificial Intelligence to visualise this difference even better.
How AI learns?
And what can we take from it
Let's imagine a random AI model, we'll call it "edux-20". It has deep knowledge of Mechanics and can solve most University level problems with ease, but now we want to make sure it gets the capability to solve Electrodynamics problems at an advanced level. Currently it exhibits medium performance on the Electrodynamics benchmarks we've created for it. How can we make edux-20, without spending money on the training process, solve those new problems at the higher level? By supplying it with external constraints - duh. We'll be uploading our Electrodynamics textbook as context with each problem we send to the model, now it cannot guess randomly since its search is heavily constrained by the contents of the text we've added (important to note that the model had possessed some basic capabilities prior to this which allowed it to interpret what the added context meant in the first place). We could represent what we're currently doing by:
(M+S)(x)→y
Where M is edux-20's neural network (equivalent of the mental model for our analogy)
S is the externally supplied context (via a Skill.md file - which is an organised way to include additional information for AI)
x is the problem on Electrodynamics
y is the answer
By doing a series of experiments
We conclude that the model has improved by 150% and thus reached (when supplied with the Skill.md file) advanced level competencies in Electrodynamics. Yet, this doesn't mean the model is now smarter.
You can take the extreme of the example, and think about using an arguably bad AI model, giving it a hard math problem, supplying it with the expected value for the solution and being astonished when it gives you back the right answer (which you supplied in the Skill.md file). You had narrowed down the search space to one single solution, therefore there was no search to be done, no conjecture to be made, no explanation to be provided, no internal change to take place. You just got the equivalent of cheating.
Naturally, the question is:
how does an LLM learn then?
We'll be skipping a lot of technicalities, but the underlying principle is correct
Learning, defined broadly, is model change. In practical terms, that means changing the behaviour without having to change the inputs. If I ask you a math question to which you don't know the answer, and then I give you the correct answer on a cheatsheet, it's incorrect to state that now you've become smarter; all that has changed is merely your inputs, but the reasoning model is still roughly the same. Akin to giving an AI a skill.md file with data, I could offer a student access to chatgpt, which will to some extent collapse the search for a guess to roughly one single possibility, so that now search in the latent space isn't even necessary since we've constrained the space too much; we now see there is such a thing as too many constraints, where model change isn't even required to produce the correct outputs.
What do we mean by model change?
Model change is a persistent change in the internal structure responsible for generating predictions, explanations and conjectures, such that the same inputs can now be interpreted or acted upon differently.
How can we train an AI model?
This is the fun part!
Let's go back to our model edux-20, now we've changed the benchmarks. We currently have problems which are at least one order of magnitude more difficult, they employ different electrodynamics concepts at once, and despite us using the edux-20 model with the Skill.md file, it performs quite mediocrely. What was the problem? The new tasks were more complex and couldn't simply employ the rules we added as context, since at some point the gap between what we need the model to do and what it can do with the augmentation is too large. The information being used only at inference was constraining the search space but it wasn't being used in its entirety as knowledge itself because it wasn't fully understood by the model (by understood we mean a deeper kind of comprehension that is embedded into how the model thinks with that information). Alternatively, some of our users reported the model not being able to solve some of the basic problems too, due to starting a new chat where the skill.md wasn't being loaded for technical issues, or overloading the model with other context (and as a result not leaving enough space for the electrodynamics one).
The solution? Embedding that knowledge within the model's brain (neural network) so that the skill.md isn't even required.
The process is quite simple, we use edux-20 (with its current brain represented by M) and the skill.md to generate a lot of "good outputs" in the following way:
Those input-output pairs become training examples:
The important part is that the Skill.md isn't included in the inputs of the training data anymore. It was used to create the examples.
The training process involves providing edux-20 with only the crude input, asking it to generate an output, and tweaking its weights (the "wring" of its brain/internal model) so that it produces better outputs each time. We compare these generated outputs against the high-quality ones produced by the enhanced version of edux-20, and keep iterating and letting our training model see how small tweaks in its internal model affect the outcome.
The moment we finish the training process, we obtain edux-21, which can produce the correct answer without being supplied with the additional context . We now have evidence that some capability previously supplied externally by has become embedded in , because the model can reproduce and generalize the behavior without being given .
We now have evidence the model has learned
Knowledge cannot be transmitted
Else we would have been done by now...
Imagine how easy it would have been if we could simply upload each others' knowledge into each of our brains and navigate the world with massively greater understanding. That might happen if Neuralink or some other startup keeps progressing and no big catastrophe happens in the next couple of decades, but it's not the case yet, so we have to deal with whatever tools and methods we have available at the moment.
It turns out that how knowledge is communicated between people is closer to how knowledge gets created, since there's roughly the same underlying logic of knowledge growth here. It doesn't matter if we, as a society, have already amassed huge amounts of knowledge in the field of chemistry, let's say... for an undergrad student who is wrestling with the topic for the first time, they still have to rediscover it all. By rediscover, I don't mean reproducing humanity's historical path to chemistry from scratch. I mean that the learner still has to reconstruct the knowledge within their own mental model; civilization's discoveries can massively constrain that search, but it cannot perform the reconstruction on their behalf. This part triggers most people from academia who are somewhat against Constructionism, to which we'll refer in a bit. This concept of rediscovering it all feels strange for two reasons:
1. It feels like a rather slow process that's absurdly inefficient. "What do you mean let people discover the knowledge themselves?!"
2. It gives more autonomy to the learner, thus taking away control from the tutor, because the methodology can no longer be built around simply transmitting information in the correct order. You have to design it around the learner doing a substantial part of the work themselves: constructing, testing, connecting, and reconstructing the knowledge.
Let's look back at how knowledge grows based on Popper's philosophy: we start with some existing model of the world, encounter a problem that this model cannot explain, generate conjectures that might solve it, and then expose those conjectures to criticism until one of them survives well enough to become part of our updated model. In a very compressed form:
Where is our existing knowledge, is the problem, are the conjectures we generate, and is the explanation that currently survives criticism better than its alternatives. This crude representation will be useful later when we compare how learning in a classroom usually happens
To refresh what this looks like for a learner: First, they encounter a problem (which means there's some discrepancy between what their model can make happen / predict and what reality actually does—in other words, their model breaks). Then they have to navigate an infinity of guesses, which they heavily filter by using the constraints given by the problem and existing knowledge, before constructing a series of conjectures which they then go on to criticise, so that in the end the conjecture which survives their critique and serves them better than the previous explanation gets adopted.
What the tasks of a learner are:
1. Understand and construct the rough shape of the problem: what it entails and what the known unknown is. If the gap is still an unknown unknown to the learner, it is not yet a "problem" in the epistemically useful sense. A problem has to become visible before it can be grappled with.
2. Search the latent space using the existing constraints (that includes the problem, existing knowledge and externally supplied information) to reach for a few testable Conjectures.
3. Criticise his conjectures until one that bridges the gap is found and embedded within the existing mental model. The embedding happens through repeatedly constructing with, testing and criticising the conjecture across different contexts, forcing it to interact with the learner's existing ideas until it becomes woven into the broader mental model rather than remaining an isolated piece of information.
In the case of education, what we have is roughly , where a more experienced human tries to transfer a compressed region of their Latent Space (call it their understanding of a specific subject) via speech, textbook, or writing to human . Human has some internal construction . To communicate it, they have to serialize part of it into an external representation : speech, writing, equations, drawings, examples. Human never receives ; they receive , from which they construct . What gets communicated are therefore constraints on reconstruction, not the original construction itself. We humans cannot communicate through telepathy; if we could, maybe transmitting higher-level representations of our thoughts would have been possible. Instead we compress and then regenerate information (due to the lack of an unified decompression algorithm). What we do transmit are constraints that can guide the learner in constructing the knowledge in a more accurate way. As such we can see why rediscovering the knowledge, to some extent, is inevitable. One cannot just upload hundreds of gigabytes of text on math and understand it. Now that doesn't mean we should rediscover it in the way it was found chronologically, or without external guidance obviously.
To make a parallel to our previous section on how AI models learn, you cannot simply transfer an LLM's "mental model" into another's unless you copy the brain (weights). The way it's done is by having the dumber LLM generate outputs that it verifies against the "better outputs" from the teacher model. If you merely supply the LLM with context, it won't learn; you will be merely enhancing its capabilities in the short term.
Humans are strangely similar in some regard. If you supply a student with the right context during a lesson and make sure they pay attention, you might end up enhancing their abilities for a short period of time. You can train them to recognise certain types of problems in physics, for example, and tell them which formulas to use, but that rarely translates into learning, because learning doesn't entail giving them information about how to solve the problem. Learning means their model goes through a lasting change that allows them to solve the problem without being supplied the full external context. It means they can search the space of conjectures better, and that they can see what the knowledge they've been supplied with entails beyond the first-order effects. That's the reason why some students can solve all the simple physics problems but fail to do well on the complex ones, when in truth the complex ones are merely a rearrangement of the simple ideas merged together. It's as if the student has been supplied with a Skill.md file that they can access for a period of time, but didn't yet fully embed the knowledge into how the reason about the problem space.
In order for learning to happen (in the context we're discussing at school), the student has to do more than successfully reason while the teacher's explanation is present as an external constraint. They have to repeatedly generate with it: use it to explain, predict, construct, compare and criticise across different contexts. By doing so, the newly constructed knowledge is forced to interact with the learner's existing mental model, exposing contradictions and creating connections until some of the search that previously depended on the external explanation becomes compressed into the learner's own model. Otherwise, the explanation can remain little more than a constraint used during inference.
Mathetics
What the learner must do with the information
When we speak about education and the educational system as a whole, what we usually think of is "educating/teaching students". Do consider how, in most instances, the student is the object of the sentence, and rarely its subject. This might seem to some like a mere grammatical triviality, but it reveals something deeper about the causal model we implicitly hold of education: the teacher teaches, the school educates, and the student is educated. Learning itself almost disappears from the picture as something the learner does.
Yet, from everything we have established so far, teaching and education are downstream of learning. They can shape its environment, supply constraints and expose problems, but the actual reconstruction of knowledge still has to happen within the learner. Teaching is therefore less about transferring knowledge and more about supplying and arranging the constraints under which the learner can construct it. The distinction between teacher and learner also begins to shrink, as functions that initially had to be supplied externally are now progressively internalised by the learner.
What the process of learning in school looks like today is closer to:
with the focus placed on pedagogy as the core way of improving the system. But improving merely half of the equation does not necessarily make education better, especially when improvements in pedagogy can only really be judged by whether they improve what we are portraying here as mathetics — the equivalent of pedagogy for learning, or the art of learning, a term used by Seymour Papert to describe the study of how learners themselves learn.
We can observe how, for one, the construction is individual-based, since it uses the student's pre-existing latent space to construct some internal representation from the information using their pre-existing model. One thing to consider is that this construction happens regardless of whether the student goes on to create an external artifact of their own. The question of "How accurate is the generation the student came up with?" cannot be answered without forcing the model to "build something", in other words to use the newly compressed generative space. The learner has to make their construction generate consequences. Assessing the new model needs to happen not only externally for the sake of a tutor, but also internally, so that the student can become more self-sustained.
These ideas converge with a whole learning theory called "Constructionism" which was developed by Seymour Papert. Constructionism focuses on how learning can be strengthened when the learner does not merely construct knowledge internally, but externalises that construction by making something in the world: a program, a drawing, a model, an explanation, a machine, a theory, or any other artifact that can embody and expose parts of their thinking. The artifact then becomes an object to think with: something the learner can manipulate, test, observe, criticise and reconstruct.
The importance of the artifact is therefore not simply that "learning by doing" is more engaging. It gives the learner's otherwise hidden construction causal consequences outside their head. A conjecture that only exists internally can remain vague and contradictory for a surprisingly long time unless it is exposed to criticism. Constructionism can be seen as a way of making the hidden process of conjecture and criticism partially observable, since construction creates unusually fertile conditions for criticism by forcing otherwise internal conjectures to generate consequences in the world. I am referring here mostly to the kind of conjectures that do not necessarily happen consciously or require much deliberate mental effort. Constructionism proposes that learning is especially powerful when learners construct public, manipulable objects that embody their thinking. In the framework we have developed here, this also makes construction useful for education because it gives both the learner and the tutor better evidence about the learner's generative model, allowing subsequent constraints to be chosen more intelligently.
To give only a few examples from constructionism, take Papert's Logo programming language. What it aimed to do was help students learn basic mathematical and geometric ideas by giving them a programmable "turtle" whose movements they could control. Rather than learning about angles from using a protractor, students were asked to make the turtle navigate a rough terrain to reach a given destination. They would do so by entering different commands into the console such as: FORWARD 10, LEFT 90, PENUP, REPEAT 4 [FORWARD 100 RIGHT 90], etc.
If one wanted to make a circle, they'd only have to input: "REPEAT 360 [FORWARD 1 RIGHT 1]". A square would be given by "REPEAT 4 [FORWARD 10 RIGHT 90]". As students were experimenting with the program, they kept updating their generative compressed model of what angles mean, they learned small subtleties about angles that they might have not understood from the initial constraints alone. In such an example, the student was learning to debug their own model through the various iterations, empowering them to be more than just a passive receiver of information.
Redefining education
The future
We have enough context by now to understand the pre-requisites of learning, what it entails and the mechanism through which learning can be accelerated.
We have, in other words, accumulated enough constraints to guess what a SOTA Educational System would even look like.
Let's review the constraints first:
1. Education is downstream of learning
2. Self directed knowledge growth begins with a problem.
3. New knowledge is constructed through conjecture under constraints. We use instruction to help shape the search space, not to eliminate the search.
4. External performance is not the same as model change.
5. Knowledge cannot simply be transmitted. What one person communicates is an external representation from which the learner has to construct their own internal version.
6. Construction makes the learner’s model visible. By explaining, building, criticising, etc. the learner exposes gaps and forces the new information to interact with what they already know.
So whatever system we come up with, it should make sure that the curriculum is built around problems whose rough shape the student can understand, problems where the student sees known unknowns rather than unknown unknowns (gaps they cannot yet even frame) or known knowns (problems their existing model can already solve without meaningful reconstruction); this allows for conjecture to take place, gives a clear target for what the known unknown is and what a successful outcome would look like, and offers enough initial constraints that are familiar to the student. The external agent (teacher, textbook, AI...) must have as good a representation as possible of what already possesses, such that their instruction can be interpreted by the student and offer the right set of constraints the learner's model was missing, or expose where that model is wrong. The environment or method through which the student completes the problem should allow for as many useful iterations as possible, without punishing the student for failing, but instead allowing them to improve across a large enough sample of attempts. The agent would then look at this progression and, based on the space the subject is searching across, if needed, supply them with more constraints or reduce parts of the problem to simpler ones until a strong enough construction around the needed space arises to let the learner continue the search with less external support, while increasingly testing whether the newly constructed knowledge transfers across different contexts.
What's wrong with the current model?
Let's look into the modern day educational system step by step, since there's too many angles from which we could understand why its shortcomings have such negative lasting changes upon its subjects.
1. Education is downstream of learning
We can view School as part of a broader system, and just like any other one (Business part of the Economical system, Lawyer part of the Legal system...), it operates within the Educational System. Every system has a way to interpret reality and filter meaning: the businessman will think through the profitable/unprofitable lens, the lawyer through the legal/illegal, and the teacher through... whatever is Education's filter for meaning.
If education is a system, what distinctions make information meaningful inside it? This is a quite interesting parallel to make with Goodhart's law which states that when a measure becomes a target, it ceases to be a good measure. Why? Because let's say you want to optimise for learning, but you cannot yet see what the mental model of each student is, so you choose a proxy - grades. Once you set the proxy and everyone knows that you're measuring the proxy, the grade fails to be the good measure, because now everyone optimises for the grade rather than it being a secondary effect of learning itself. And this refer to students only, but also teachers, institutions, parents, etc.
The lens through which education filters meaning, is as we all know it passed/failed, it's a ranking system, and this is what all the participants within the system optimise for. In a way, you can argue that education as a system is quite brilliant and successful at what it's doing, but here's where the discrepancy comes from when people wonder why we cannot innovate the system in any shape or form; we're trying to push apps and methods that optimise for learning in a system that doesn't, so it rightfully kills anything else...
2. Self directed knowledge growth begins with a problem
There's no reason to change any object's state if it's in equilibrium... Same with a human's mental model, as long as you don't criticise it or get it exposed to problems where the model breaks, there's no reason for it to change... Merely introducing new tools without explaining why we need them will result in disengagement. And it also teaches a more harmful lesson than people realise, which is using the wrong epistemic order to create things.
First of all, not starting with a problem makes the learning much harder because you don't really have anything clear to optimise for, and optimising for everything means optimising for nothing.. As a result, you can come up with complex structures of conjectures which in the end amount to a bunch of nonsense. I had to unlearn this lesson myself as I got into startups, and you see it all the time... smart people building things nobody wants because they never had a clear or somewhat accurate feedback loop on whether or not they were progressing or regressing...
In most schools however, problems are a mere feature you do after the lesson has been taught. This is the wrong way to think about it because now you failed to let the learner understand the shape of the problem, which would have supplied a lot of the constraints...
3. New knowledge is constructed through conjecture under constraints. We use instruction to help shape the search space, not to eliminate the search.
In traditional schooling, you are supposed to teach the material and then let students solve exercises which employ the exact same knowledge; hardly any knowledge gets constructed. Instruction is what every new lesson starts with, but this model was made to fix an underlying issue - scarcity of access to information. In the 21st century, we are lucky not to have to worry about that problem. Instruction doesn't equal knowledge transfer. Supplying novices with substantial knowledge in a field they have no experience in can make a tremendous difference, because obviously, they had no other option to do inference on a given problem outside us supplying them with scaffolding. But the principle is that what we call knowledge transfer is in practice knowledge creation at the individual level with external support... Someone inexperienced benefits from more external constraints while someone already experienced might benefit from less... It's never about the instructions, but always about the construction that happens at the individual level.
4. External performance is not the same as model change.
Most tests have no way of telling whether the people they're assessing have internalised the knowledge or are merely using it as scaffolding. While we cannot yet assess model change, we could develop better ways to reduce the likelihood that what we're assessing is merely how well has the student retained the instruction . [2]
5. Knowledge cannot simply be transmitted. What one person communicates is an external representation from which the learner has to construct their own internal version.
The way schools teach pupils is by continuously telling them what is the right way of doing things, what is out there and what you must take as a fact... The reality is that the so-called "right way" is coherent within the lesson's parameters (otherwise it wouldn't have survived criticism) and a lot of the times will "make sense" for a lot of students. But the core reason for which instructions exist is to supply constraints. When the student is disconnected from the lessons, the constraints entailed by the instruction seem unintelligible for him... As a result, the student would benefit more if he were explained why "the wrong way" of doing things was wrong... That usually would happen by having him make a tiny prediction and test his mental model against reality. As such he would start the process of criticism and conjecture-building that would allow him to understand the new material.
6. Construction makes the learner’s model visible. By explaining, building, criticising, etc. the learner exposes gaps and forces the new information to interact with what they already know.
The primary way the student's model is criticised is usually through formal examination, and even when teachers do hold discussions in the classroom, formal examination is still what ultimately filters students and decides how much “learning” has supposedly occurred. By formal examination I mean: essays, tests and, occasionally, oral examinations. But we rarely see students taking on substantial projects where their models have to survive contact with real constraints and produce real solutions. And I'm not talking about masquerading around with Project Based Learning initiatives that are often no closer to genuine projects than education itself is to learning. [3]
This is, a complete understatement to how deplorable the situation is... But I think it paints the picture well enough, and allows anyone who has followed carefully the arguments the ability to see the other issues we haven't spoken of, themselves. By now you hopefully got a good enough sense of what learning is, in order to spot everything else that it isn't.
What does the Future of Education look like?
In the following essay, I will go further down the rabbit hole to see what would a State-Of-The-Art School look like, and see how wide is the gap and whether it can already be filled now with the current technological advancements.
Endnotes
Endnote 1
Most people’s complaints about school concern how inefficiently it equips students with useful knowledge and capabilities. Whatever other functions schools might fulfil, their declared purpose is, first and foremost, to give pupils the essential knowledge and skills they need. I am therefore evaluating schooling according to the objective it publicly claims to pursue. If education is not actually its primary function, that would not weaken the argument; it would help explain why the system can remain successful on its own terms while repeatedly failing at learning.
Endnote 2
The best way to assess whether someone understands a topic is to see whether they can reconstruct what it entails by making something with it. Richard Feynman famously wrote, “What I cannot create, I do not understand.” The idea is to force the person to expose how they think about a specific topic from multiple angles, rather than merely asking them to reproduce the same information in roughly the same form in which it was taught.
Creating something is especially useful because it makes this harder to fake it. It forces the learner to expose whether the knowledge has actually become generative within their model, or is it merely remaining something they can retrieve and apply during inference.
Endnote 3
Constructionism isn't precisely about giving students a project to complete, but about exposing their mental model to themselves so they can criticise and improve it by going through enough feedback loops.
What is usually presented as PBL, however, is closer to a lesson where the teacher assigns a project to the whole classroom, or to groups of pupils, with the expected direction and outcome already largely constrained. The project then becomes more or less distributed homework which the students divide among themselves. As a result, it becomes difficult to see what any individual student actually understands, and, more importantly, each student can complete their part without necessarily having to make their own mental model generate anything, encounter failure, or reconstruct itself. The final artifact may look like a project, but epistemically it often amounts to little more than regurgitated information arranged into a presentation, poster or report.
