Skip to content
Ben Stewart

posts

What Level Does This Task Deserve?

Two questions I ask before I hand anything to AI

· 5 min read

A single trading card reading Level 5, range 3 to 7, beside a hand of five task cards: thinking work at Level 3, finishing a pattern at 4, anything an exec reads at 5, mechanical fan-out at 7, and a face-down Level 6 card marked switched off.

Somebody printed their AI level on a trading card and I haven't stopped thinking about it. Jaatster's card says LEVEL 8 ORCHESTRATOR in big pixel type, and it's a great piece of work. It's also where this post came from, so credit first; here's the tweet.

The bit I keep coming back to is the footer. In much smaller type it says "Workflow foundation: Level 5". Jaatster printed both numbers, which is more honest than most of the frameworks behind these cards, because a single number on a card has to be the best thing you've ever done with the tools. Nobody prints their median.

The cards come from Every's Eight Levels of AI Adoption, a prompt that asks you five questions and hands you a level. I took it with one change. I didn't answer the questions; I told the AI to go and read a month of my own session transcripts and grade me on what I'd actually done.

Seven of those sessions were substantial, and the biggest had 22 prompts from me. It came back with Level 5, confidence medium-high.

sessions in a month
24
tool calls in one session
388
subagents in one go
37
my level (3 to 7)
5

I'd have scored myself higher. Everyone would. You answer five questions thinking about your best day and call it your level.

The range was the interesting bit, though. When I looked at where the 3s and the 7s came from, they followed the work, and if I'd been handed one number at the end of it I'd have had to pick which Tuesday to be.

Why everyone wants a level

I get the appeal. A level turns "am I behind?" into "I'm a 4", which is a much nicer thing to carry around. The prompt ends in a subscription link, so there's always a next rung for sale. A number also fits on a slide, and anything that fits on a slide gets counted, and anything that gets counted eventually becomes a target.

The idea is borrowed from self-driving cars, where levels describe a machine. Point it at a person and it sounds rigorous and isn't. Every framework I've read says higher isn't better, then lists the signs you're ready for the next level, which tells you which way they expect you to go.

I don't want to be sour about this, because the levels gave me something. I can now say "that was a 3 job" or "that one needed a 5" and mean something specific, and I couldn't do that a month ago. I'd keep that. I'd drop the idea that the number describes me.

Two questions

Before I hand a task to AI now, I ask two things.

What does it cost if the output is wrong? Can I check what comes back?

Put those on two axes and you get a grid.

Expensive if wrong · Easy to check

Level5

Let it work, then check everything

Anything an exec is going to read

Expensive if wrong · Hard to check

Level3

Stay close

Thinking work

Cheap if wrong · Easy to check

Level4 / 7

Let go, check a sample

Finishing a pattern, mechanical fan-out

Cheap if wrong · Hard to check

Level6

Ask whether it should run at all

The weekly research sweep, switched off in July

My month fits on it more neatly than I'd like to admit.

Thinking work sits top right. Working out what I actually believe about something is expensive to get wrong, because I end up holding a weaker idea and presenting it as mine, and there's no answer key. My instruction there is nearly always some version of "just tell me what you would do, don't do it". That's the 3 on my card, on purpose. Some of it shouldn't go near the tool at all yet, and no ladder has a rung for that.

Anything an exec is going to read is expensive too, though you can check it, because numbers trace back to sources. So the AI does a lot of the work and I say "I need every single number and claim checked" before any of it goes in front of senior people. That's where the 5 came from. The AI does more here than on the thinking work and I trust it less, which still feels slightly backwards to me.

Bottom left is where I let go. Once I've done the first few of something and know what good looks like, I say "do everything that's left" and stop watching. That's a 4. The 37 subagents live here too, a manager model handing a pile of repetitive linking work out to cheaper models in one session. I didn't read 37 outputs. The highest number on my card, the 7, came from the most boring work I did all month.

Bottom right is cheap and hard to check. I used to run a weekly automated research sweep, and nothing it found was worth reading without me asking for it first. I switched it off in July. There are no scheduled AI jobs on my machine now, so Level 6 on my card is a zero, and I'm fine with that.

Every ladder I've seen only goes up. The best line I read on this all month was from u/Fyren-1131 on r/ExperiencedDevs.

I am trying to find the level of AI use that makes sense for me.
u/Fyren-1131, r/ExperiencedDevs

Somebody with a Level 8 card and somebody switching things off can both be doing that well.

Try it on Monday

Write down five or six things you handed to AI last week and put each one in a box. You're looking for two kinds of mismatch.

Something expensive and hard to check that you left to run on its own is the one that burns you, usually weeks later and in front of someone.

Something cheap and checkable that you're still reading line by line is where your week went. You've built a very fast intern and then stood behind it all day.

Then look at the bottom right box. If something's running there, ask what you'd lose by switching it off. For me it was nothing.

As a senior leader I'm constantly thinking about what this means for how people work today and in the future. If I could change one thing about level frameworks, Every's included, I'd score the task and leave the person out of it. A person's level is just the mix of tasks on their desk this week. Mine moved between 3 and 7 inside a month and I didn't get any better or worse at this in between.

TL;DR

I asked AI to grade my AI level from a month of my own session transcripts. It came back as Level 5 with a range of 3 to 7, and where I landed each time depended on the task in front of me. So I've stopped asking what level I am and ask two questions of every task instead. What does it cost if the output is wrong, and can I check what comes back? Score the task and leave the person out of it.


How this was written: drafted with Claude, rewritten by me. Fitting, given the subject. Claude read the month of transcripts, three different models drafted versions of this post, two blind judges scored them, and I rewrote the winner by hand.