Skip to content
All Articles

I Don't Explain My Pedagogy Anymore. I Wrote It Down Once.

19 min read
aieducationmcpskillsedtechworkflow
Share
I Don't Explain My Pedagogy Anymore. I Wrote It Down Once.

The other day I typed a sentence shorter than any email I wrote that evening: "Check the exercises in the bookkeeping module for me."

What came back wasn't an opinion. It was a table. 26 exercises, each scored against five criteria, the four weakest at the top, with reasons why they were weak.

The point isn't the speed. The point is that I did not use a single word that evening to explain what makes a good exercise. I had written those five criteria down months earlier, exactly once. They have been sitting there ever since, and the AI fetches them itself the moment a question sounds like it needs them.

People often ask what is special about my setup. The expected answer is a number: 18 connected services, some forty skills, a Moodle server with more than ninety tools. But the number isn't the answer. The answer is an interlocking of two halves that is easy to miss, because each half sounds unremarkable on its own.

Two halves that are worth little alone

Let's start with the terms, briefly and assuming nothing.

MCP stands for Model Context Protocol. It is a standard for sockets. An AI can talk and write out of the box, but it cannot touch anything. An MCP server is a small program that sits in front of a real system, say Moodle, and offers the AI a list of moves: create a course, rename a section, fill a quiz with questions, read out grades. Once that server is plugged in, the AI can perform those moves in your real system, with your real permissions. MCP gives it hands.

Skills are the other half. Technically, a skill is a folder with a text file in it. At the top of the file are two lines: a name, and a description of when this skill applies. Below that, in plain language, is how the work gets done here. Which criteria apply. In what order. What separates a good result from a mediocre one. No code, no programming. The AI reads that file only when the description matches the task at hand. A skill gives it judgement.

The difference becomes clear with a single question:

Your question is essentiallyThen you need
"Reach into this system"MCP
"Do it the way it's done here"Skill
"Now actually create it in Moodle"MCP
"By what standard is this good?"Skill
"Do it like we did it last time"Skill

I built both halves one after the other, and I experienced each of them alone. Both were disappointing.

The MCP half first. I have written about that already, in 12 servers, one protocol. The feeling of typing "Create a quiz on breaches of sales contracts" for the first time and finding a finished quiz thirty seconds later was enormous. It carried me for about three weeks. Then I noticed I was producing a great deal of mediocrity very quickly. The questions were correct and boring. The courses were complete and lifeless. I had given an AI hands and forgotten to tell it what I consider good work.

The skill half alone is just as unsatisfying, only the other way round. An AI that knows your quality model but can't reach anything hands you an excellently reasoned recommendation. Which you then work through yourself. On Sunday evening. By hand.

Practice 1: Build for the systems you actually work in

Not the ones everybody talks about. My first MCP server could do three things in Moodle, because Moodle is where I sit. Anyone who starts by connecting a system they touch twice a year is building a showpiece, not a tool.

The second time, it's a skill

The obvious question: how do I know when something should become a skill?

My threshold is lower than most people expect. It sits at the second time.

When I explain the same thing to an AI a second time, that is no longer a coincidence, it's a pattern. I am not explaining twice because the model is forgetful, but because I hold a requirement so self-evident to me that I have never said it out loud. Those requirements are precisely the valuable ones. They are the difference between my work and anybody's work.

The first time, explaining is normal. The second time, explaining is a negligence that compounds.

In practice that means: I no longer put the explanation into the chat, I put it into a file. The second time round that costs about four minutes more than explaining. The third time it breaks even. From the fourth on it's profit, and permanently so, because the file doesn't get tired and doesn't go on holiday.

Practice 2: The second time you explain it, it's a skill

Not the fifth time, not "when I get around to it". The second time. What you explain twice, you will explain twenty times, and every one of those is lost time plus the risk of explaining it differently next time.

And then comes the part I underestimated at first: a skill is useless if it isn't found. The AI doesn't read forty files before answering. It reads the description lines and decides from those which file gets opened at all.

Which means: the description isn't the packaging. It is the tool. An excellent skill with a vague description is a book in a library with no catalogue. So my descriptions contain the words I actually use when I'm tired: "worksheet", "handout", "create course", "end of session", "we're done". Not the words that would appear in documentation.

Practice 3: The description is the real work on a skill

Write down when it should apply, in exactly the words you use day to day. A skill that isn't found doesn't exist. And the more skills you have, the more that one line decides the worth of all the others.

Why only the product of both carries weight

You can picture the situation as four fields, and all four occur in daily life.

A two-by-two matrix of reach and judgement: neither, judgement only, hands only, both.
The dangerous corner is bottom right, because it feels like productivity.

Neither. You describe your problem to the AI, it answers cleverly, and afterwards you do everything yourself. That is the normal case for most people, and it is not nothing. But it is advice, not work.

Hands only. The AI can reach everywhere and has no idea what is good. It then produces a great deal of material very quickly, material that passes muster and reaches nobody. This is the most dangerous corner, because it feels like productivity. You see the volume and miss the fact that quality didn't grow with it.

Judgement only. The AI knows your standard and can apply it nowhere. The result is a good plan. Executing it is on you.

Both. And here something happens that doesn't feel like addition. Because the judgement is permanently on file, the character of my instructions changes. I no longer say what is to be done. I say what I want to achieve. "Create a quiz with eight questions on breaches of sales contracts, two of them multiple choice, make the wording action-oriented, no pure recall questions" becomes "Build me the opener for unit 3."

The sentence got shorter. It did not get vaguer, because the precision moved house. It now lives in files instead of my short-term memory.

That is the actual return, and it isn't a time saving. It is a shift: away from instructing, towards deciding.

Four principles from my setup

I'm not describing my skills here. They wouldn't help you much, because they contain my subjects, my school and my corporate design. I'm describing the four ideas behind them, because those transfer.

A skill carries a model, not a sequence of steps

The skill that graded those 26 exercises contains no instructions. It contains a model with five criteria: is there a story anchor the exercise hangs from? Is there anything to see? How deep is the thinking being asked for? What is the student actually doing, apart from reading? And: what happens when they get it wrong, is the error punished or explained?

Five questions. No procedure, no order, no technical instruction.

The difference is bigger than it looks. A sequence of steps ages with the tool it was written for. A model survives the change of tools, because it describes how you recognise good work, not which button to press. My course analysis works from a four-competency model, my exercise check from those five criteria, my school documents from the school's corporate design. In all three cases I wrote down a professional conviction, not a manual.

And this is where teachers have a head start they don't know about. These models are already in their heads. They have simply never had a format in which they could be passed on.

Practice 4: Write down the model, not the steps

How do you recognise good work? That question is the substance. The sequence of clicks is not, and it changes with the next update. A skill carrying a quality model will still be valid in three years.

Measuring and improving are two separate tools

I work in pairs. One skill that analyses course sections, and a second that optimises them. One that grades exercises, and one that reworks them. Always two, never one.

That is deliberate, and it cost me something to arrive at. A tool that measures and repairs at the same time flatters its own measurements. It has an interest in the outcome. It finds the problems it can solve and reliably overlooks the ones it has no answer for. What you get in the end is a report in which everything looks fine, because everything bad was quietly fixed along the way and nobody can remember how bad it was to begin with.

Separate tools force an intermediate stop: I see the measurement before anything is changed. I can object. Sometimes the weakest exercise is exactly the one that works best in class, because it has a story no criterion captures. That decision belongs to me, and it only belongs to me if I see the numbers before the intervention.

Practice 5: Separate measuring from improving

Two tools, two calls, one decision in between. Do both in one step and you get flattering measurements and lose the moment where you could have objected.

A skill conducts, it doesn't reach in itself

My largest skill builds complete teaching units. It researches sources, plans the pedagogy, produces material, checks its quality and finally publishes to several platforms. What it does itself is: none of that.

It describes an order and hands each step to the hands that can do it. Research to the research route, material to media generation, publication to the Moodle connection and the learning module portal. The skill is a conductor, not a musician.

That sounds like a nicety, but it is the reason the thing has been running for months. If a single connection changes, a single connection changes. The sequence stays valid. Had I written the execution into the skill, I would have to touch it every time any participating system changed.

One skill in the centre, arrows out to research, media, Moodle and portal.
The skill holds the sequence. Each step is carried out by the connection that can do it.

The same reach, different identities

This is the part of my setup I explain least often and consider most important.

My connections don't run straight to the servers, but through a thin intermediate layer. It does one thing only: it decides which key gets used, depending on the context I am sitting in. School is one context. The blog is another. Server operations is a third. The association I volunteer for is a fourth, and it is strictly separated from school.

From the AI session through a context layer to four separated contexts: school, blog, server operations, association.
Four contexts, four keys. The separation lives in the plumbing, not in self-discipline.

The effect is twofold. First, I can see afterwards in a log what happened in which context. Second, and this is the real gain, a session meant for the blog cannot accidentally reach student data. Not because I am disciplined, but because the key for it does not exist in that session.

The same logic applies to the most dangerous mix-up of all. I run two Moodle installations: a local sandbox and the real one, with real courses and real students. They are deliberately named differently, and the rule is written down: if it isn't unambiguous which one is meant, ask, don't guess.

Practice 6: Give the AI several identities, not one omnipotent one

Separate the contexts you work in and give each its own access. The best protection against a slip isn't caution, it's an access route that simply isn't present in this session.

Practice 7: Two identically named systems need two names and a rule

Test and production, old and new, drill and the real thing. Name them differently and write down the rule for the case of doubt. Mix-ups don't come from stupidity, they come from similarity under time pressure.

The working habits around it

So far this has been about tools. But the part that changed the most for me isn't the tools, it's three habits around them. They cost nothing and work without a single server.

The first: every project gets a folder saying what was done here and why. For me those are dated design documents, implementation plans and reports on what actually came out at the end. My infrastructure project now holds sixteen such designs.

The reason isn't tidiness. The reason is that a chat window dies.

An AI session has a memory that ends with it. Everything you worked out together over three hours, every dead end, every trade-off, every reason you settled on the second-best route in the end, is gone afterwards. Come back in six weeks and you are facing someone seeing your project for the first time. And you yourself no longer remember the reason either, only the decision.

A project folder is the answer to that, and it has an addressee you don't immediately think of: the next agent. I don't write these documents for myself. I write them so the next session reads them and reaches in five minutes the state that took us three hours last time. That is exactly the point where individual good evenings turn into something cumulative. My work doesn't grow because I type faster, but because each session builds on what the previous one left behind.

Practice 8: Give the project a docs folder the next agent will read

Not for posterity, for the next session. What you decided, why, and what you discarded. Without that folder you start from scratch every few weeks and don't even notice.

The second habit: everything lives in version control. For me that's Git, mirrored to GitHub, and it applies to the code as much as to the teaching materials, the skills and the documentation.

This usually gets sold as a tidiness topic, which is why hardly anyone listens. But it isn't a tidiness topic. It is the precondition for being able to work boldly at all.

The order runs like this: safety first, freedom follows. The only reason I let an AI work unsupervised across forty files at once is that a single command brings back the state from a minute ago. Every change is a proposal, not an intervention. Without that return ticket I would get careful, and caution is exactly what destroys this way of working. You then only try what you already trusted yourself to do, and that's the barrier I wrote about in Three barriers that no longer exist.

There is something else that only matters later: version control makes work inheritable. In I rebuilt it I described why public authorities shy away from home-built solutions, and the objection is a fair one. A repository someone can clone and hand to an AI to read is half the answer to it. A folder on your desk is no answer at all.

Practice 9: Version control isn't tidiness, it's freedom from fear

Safety first, freedom follows. Someone who can always go back allows bigger interventions and learns faster. And what is versioned can be taken over by somebody else when you're no longer there.

The third habit is the least popular: I record what I measured, with a date, and I also record what turned out to be wrong.

My project documentation contains sentences like: "Measured on the live system on 16 Aug, the earlier note on this was not reproducible." It looks like bureaucracy. But it is the thing that most reliably protects me from duplicated work. A refuted assumption written down nowhere comes back. It comes back guaranteed, in four months, when nobody remembers it was already checked, and then it gets checked again.

That goes double for AI sessions, because an AI is happy to agree with you. It will adopt your old wrong assumption and build on it, if that's all your notes contain and nothing stands beside it.

Practice 10: Record what you measured, with the date and with the error

Not just the result, but also the assumption that turned out to be wrong. A refutation noted nowhere comes back and costs you the same hour a second time.

Why this works better in schools than elsewhere

Now the part I actually care about.

I'm often told my setup is an exception because I'm an administrator and run servers. That holds for the number. For the core it doesn't, for a reason that almost never comes up in the debate about AI in schools.

Teachers already own the harder half. Not roughly, but completely, and at a depth that is bought expensively elsewhere. A colleague who has taught English for fifteen years holds a complete internal model of what separates a good task from a bad one, where a class typically veers off at that point, which mistakes are productive and which merely frustrate, how to pitch the same thing at three different levels. They have never written it down, because there was never an addressee. You can tell a trainee teacher, and that's it.

Now there is an addressee. And the translation work isn't programming, it's one sentence: write down in plain language what you would explain to a trainee.

The other half, the reach, is likewise already present in schools, just unnoticed. There is a learning platform, a mail system, class lists, an office suite, a filing system. Those are exactly the systems you can attach hands to. The effort is real, but it is one-off and shareable. And it doesn't have to come first.

The honest addition belongs here: the smallest sensible setup consists of zero servers. A single skill, a text file with your quality standard, with no connections at all, already gives you the effect from the first section. Your instructions get shorter, your results more consistent. You can add the hands years later, or never.

The honest price

There are three things that don't work, and texts like this one tend to keep quiet about them.

Skills rot. A skill pointing at an interface that has changed is worse than no skill, because it creates false confidence. A few of my descriptions are by now older than the systems they refer to. That costs cleanup time regularly, and I am not good at it.

Too many skills cancel each other out. At some forty of them, only the description line decides what gets found. Two skills with similar descriptions are worse than one, because then whichever happens to sound better gets pulled. The collection needs tending or it runs wild.

Speed doesn't bring quality with it. That is the lesson that cost me the most. I had to build the quality check explicitly, as its own tool, because speed on its own only generates more material. Anyone who believes good results appear once the friction is gone will produce quiet mediocrity in large quantities.

And the fourth thing, which isn't a price but a limit: all of this still hangs on one person. I have tools, not an institution. What I build outlives me only because it is versioned and documented. That is better than nothing and less than a succession plan.

The first step costs twenty minutes

If this has reached you and you don't know where to start: not with a server. Not with a connection. Not with a list of tools.

Start with the file.

Take the one thing you explained for the second time this week. Write down in plain language how you tell it has been done well. Not how it's done, but how the good version is recognised. Write one line about when this should apply, in the words you use when you're tired. Put the file in a folder your AI reads.

That's all. Twenty minutes, no server, no budget, no permission needed.

And when you notice, the time after next, that you no longer have to type that explanation, you will have built the half that is harder to come by. The hands are technology. The judgement was you.

Dirk Schulenburg, Hamburg. Writes things down so he doesn't have to say them twice.

About this text: the outline, the first draft and the examples from my own setup came out of a conversation with Claude. The thesis, the selection of the ten practices and the final edit are mine.

Related Articles

Ich hab's nachgebaut. Sie nahmen trotzdem das Original.

Die Schulzentrale kündigt TaskCards. Ich baue in fünf Tagen eine barrierefreie, self-hostbare Alternative, zeige sie, und sie nehmen trotzdem das Original. Über den Macher-Reflex, seine Grenze und eine Schranke, die gerade verschwindet.

aieducationedtechself-hostingstigmergy
5 min read
Three Barriers That No Longer Exist

78% of commercial job tasks are automatable. 184 hours of development per hour of e-learning. And we only attempt things we believe we’re capable of. Three barriers that AI is blowing apart right now, and what it means for education.

aieducationstigmergypost-workedtech
12 min read
I'm Automating My Own Job, and Nobody Notices

12 MCP servers, 73 Moodle tools, 16 H5P content types: How a teacher automated his entire content creation, and why the school system doesn't even notice.

aieducationautomationmcp
5 min read

Comments