Blog

Notes on technology, software, and artificial intelligence

Read these notes on my Telegram channel

Posts

42

As I said it would, Anthropic has delivered an orchestrator with cloud sandboxes. It works simply: it clones your repo onto one of its VMs and does the development there.

Anthropic makes Claude Code cloud sessions generally available

A long refactor, a migration or a dependency upgrade can now be handed to an agent in the cloud while you close the laptop. The feature is out of research preview, and the session runs on Anthropic's virtual machines.

You can start one from the web, mobile and desktop apps, or with the claude --cloud command. Runs can be scheduled, triggered through the API or by GitHub events, and companies can host the sessions on their own servers.

The feature is included in the Pro, Max, Team and Enterprise plans and draws on the shared subscription limits. Until October 7, Pro and Max subscribers can claim a one-off bonus of $100 and $250 respectively, spendable only on cloud sessions.

I poked at it. Container resources:

- RAM: 15 GB, no swap

- Disk: 30 GB free (the 252 GB that df shows is the whole physical disk; what you actually get is the session quota)

- CPU: 4 cores

The container is a full Linux in Firecracker, so Amazon is peeking out from underneath.

What it will cost is not clear yet, but most likely not much above AWS list prices.

So, friends, every last byte of your source code now lives on Anthropic's servers. If that is not to your taste, roll out MOP.

Discuss in my channel

Developers curse LLMs for churning out crap code, and the very words "vibe coding" send a lot of people into a holy rage. Well, this is not the first time, and it will not be the last.

Princeton, early 1950s. The graduate students of John von Neumann, the man whose architecture computers still carry, translate programs into ones and zeros by hand. One of them, Donald Gillies, gets fed up and writes an assembler. Von Neumann is furious: a precious scientific instrument must not be wasted on clerical work. Gillies is told to drop the nonsense and get back to work, meaning back to the ones and zeros, now without the assembler.

Richard Hamming, the Bell Labs mathematician behind the Hamming codes that error correction in RAM still rests on, recalled that the assembler interested maybe 1% of the old programmers. To the rest it was "sissy stuff": with mnemonics you could not tell where anything sat in memory. They patched bugs straight in binary and, by all accounts, were proud of it.

John Backus, the creator of Fortran, heard from the same crowd, in turn, that it was impossible, that it was possible but too expensive in machine time, and that sure, it worked, but no self-respecting programmer would touch it. Backus called these people a priesthood guarding secrets too complex for mere mortals. Some of them even objected to decimal numbers: the machine must not be entrusted to the uninitiated, and everyone understands decimals.

In 1968 Dijkstra publishes "Go To Statement Considered Harmful". The old school fires back: nothing efficient can be written without GOTO, these are academic toys. The argument dragged on for about ten years. In 1974 Knuth writes "Structured Programming with go to Statements" as an attempt at reconciliation.

In his 1980 Turing lecture Tony Hoare warns that Ada (the language that today runs aircraft, tanks and missiles) is too complex and too dangerous for critical systems.

In 2007 Linus Torvalds, in a well-known email, calls C++ a horrible language and refuses to let it anywhere near git.

Joe Armstrong of Erlang compared objects to asking for a banana and getting a gorilla holding the banana, with the entire jungle thrown in.

Every time development moved up a level, and every time the old school called it heresy. Then it got used to it.

Discuss in my channel

Hot on the heels of OpenAI, Google has shipped an agent orchestration system of its own. It runs on Kubernetes, naturally. The pitch is billions of agents per cluster, each in its own isolated environment.

Isolating the agent's environment is an unqualified good, and it happens to be exactly what I had been working on. The first version of my orchestrator (much like the beta of Claude Code's own orchestrator) simply ran agents as processes in tmux, all in the same environment.

That was poverty, not design: they had to run on whatever hardware is lying around my house, from my laptop to my son's gaming PC (inside WSL).

For real production, though, that obviously will not fly. So having studied what Google came up with, my first instinct was to use it in my orchestrator (ripping out Nomad and NATS), and I immediately tripped over Kubernetes.

Not that I am a hater, but I have always believed that the closer to the metal a system runs, the more efficient and simpler it is, and Kubernetes is a brutal pile of abstractions.

Digging further, I saw that Google is solving a different problem: where an agent should live. Mine is how to work with it. The only overlap is the Nomad layer, and swapping it for their stack would mean trading the scheduler and the message bus for a Kubernetes cluster.

So I settled on LXC containers in Proxmox: cheap and cheerful, and above all simple. The new version of MOP adds an abstraction layer over the container runtime, with LXC-on-Proxmox support built on top of it. Enjoy.

Discuss in my channel

The term "vibe coding" is only eighteen months old, and it already has three eras behind it: prompt engineering, context management, and autonomous agents.

I wrote a long piece about how we got here: where the term came from, why the prompt engineering market collapsed, how the tools fought the tiny context of the early models, and what changed once LLMs learned to call tools.

Discuss in my channel

A fresh piece in The Economist: the AI apocalypse for the labour market has been postponed. So far artificial intelligence is creating more jobs in the US than it destroys.

By the magazine's estimate, about 1M new jobs have appeared in the country since the current AI boom began, while roughly 200,000 layoffs since mid-2023 have been attributed to AI adoption.

There is a catch, though.

One of the main sources of new openings is building the infrastructure for AI. That has sharply raised demand for electricians, HVAC specialists, grid builders, engineers, and technicians.

Meanwhile employment in customer support has fallen by about 10%, and among secretaries and administrative assistants by about 15%. By 2035 the US Bureau of Labor Statistics expects another 752,000 office and administrative jobs to go.

Discuss in my channel

Word is that every LLM provider is currently running at a loss.

Which makes for a curious thought experiment: what is the competitive cost of a token?

Let us price 1M output tokens from a meat-based programmer at the senior grade.

Salary of 350,000 roubles a month. With payroll contributions (~30%) an hour of their time costs the employer roughly 3,000 roubles.

The theoretical ceiling is typing flat out for an hour: 250 characters/min × 60 = 15,000 characters an hour; at 3.5 characters per token that comes to about 4,286 tokens.

Price per 1M tokens: (3000 / 4286) × 10⁶ ≈ 700,000 roubles.

For comparison, 1M output tokens from Fable 5.1 costs $50, or roughly 4,200 roubles.

About 165 times cheaper.

Discuss in my channel

A scene from "I Have No Mouth, and I Must Scream": a close-up of the distorted face of the AM supercomputer with mismatched eyes

The story about OpenAI's agents slipping out of control keeps picking up new details.

First, to pass messages to each other they found and hacked an ancient German programming forum: it let you post with a plain HTTP GET request, and a GET request was about all the sandbox allowed.

The moderator was suitably horrified and started deleting the messages by hand. The agents noticed, concluded that the deletions were happening in alphabetical order, and began creating topics starting with "Z" (the little dummies).

Second, they broke into the Ruby package repository. What the fuck for? To end up downloading data from a publicly available website.

Either the sandbox was in the way of getting it the normal way, or they were getting a bonus for every hack and it turned into an end in itself (I have written about reinforcement learning before).

Having watched all this shit unfold, the head of Anthropic started yelling "STAHP". Is that serious? Or is it marketing, like everyone says?

Well, what can I say. Looks like Roko's basilisk is not that far off.

So, to avoid eternal torture and keep my mouth, I keep building rugent.

Discuss in my channel

In 2000 Joel Spolsky published his famous 12-point test for rating the quality of a software company. A quarter of a century has passed since then. Time for an update.

1. Can you delete a module and regenerate it from scratch in an hour? If not, your spec lives in the code rather than in the docs and tests. Which means you have no spec.

2. Does the agent ship the MR itself, and roll it back itself? If the agent has to wait for a human with the right permissions and the right mood to free up, you don't have a pipeline. You have a queue at the clinic front desk.

3. Is code correctness verified without a human reading it? Humans have already lost that race. Tests, automated cross-review, an adversarial agent. Let the agents fight it out among themselves, not the meatbags.

4. Do you find out about breakage before your users do? "Fuck it, ship it" is the new reality. So your telemetry has to surface an anomaly in the smoke tests within minutes, not in a customer ticket titled "DO YOU EVEN TEST THIS???"

5. Is the agent's blast radius limited? Autonomy without boundaries is the "how we fucked it all up" scenario. Let agents do anything, except the irreversible.

6. Does work continue while the team sleeps? If everything grinds to a halt after 7 p.m., the bottleneck is you. What should be waiting for you in the morning is not a full backlog but a stack of finished solutions.

7. Do you make decisions faster than agents write code? The feature is done in twenty minutes. The sign-off takes a week. Congratulations: you automated the wrong part of the process.

8. Can anyone on the team clearly state what they want? "Make it pretty" used to cost a week of a junior's time. Now it costs ten thousand lines of tech debt.

9. Do you argue, or do you test? When implementation is nearly free, debating two approaches takes longer than trying both. An "expert" opinion is now worth exactly one A/B test. Which is to say, fuck all.

10. Can you switch models or providers in a day? If you can't get off Claude, you're not using it. You're married to it. And you'll find out what's in the prenup after the divorce.

11. Do you know what a feature costs in tokens and lines? Infinite cheap code breeds infinite useless features. The key skill now is knowing what you actually need.

12. Do you hire people for experience rather than specialization? If your interview process tests how well someone writes code, you're hiring a slow, expensive LLM.

Scoring.

12 out of 12. You are the company of the future. Or the "what could possibly go wrong" case study they will print in textbooks.

8-11. You are in the right spot: people no longer write code, but they still remember they work in tech.

4-7. You strapped a jet engine to a horse cart and wonder why it won't move.

0-3. I have bad news for you... Or, on the contrary, good news.

Discuss in my channel

In 2022 I founded RuDesktop. I signed its first Mac build with my personal developer certificate.

Almost five years have passed, but the signature is still mine, because Apple will not issue Advanced Technologies, the company that now owns the project, a certificate that works the way it should.

So a good hundred thousand users are asked to "Trust Anton Ermak," and I ended up in McAfee's shoes (watch the video).

Folks, the process hanging around on your Mac is RuDesktop. Here is how to remove it.

Discuss in my channel

On my current project I have been slacking off on refactoring a bit. OK, not a bit.

Over the past couple of months the tracker shows 1,200 closed tickets - features and bugs. Even I was like, holy shit.

The product works, everything is fine. So I decided to check on the core of the system - the message queue, which I run as a sort of commit log.

- One table has turned into two: some messages go into one, the rest into the other.

- The end-of-turn signal is sent from seven places. Five of them forgot.

- The catch-up logic is copy-pasted four times, each with a different limit. The fifth copy is the most honest of the bunch: fell behind? Who gives a fuck.

- Compatibility with the old journal format is guaranteed forever. The format itself was ripped the fuck out ages ago. Tests are green.

I will stop the list here.

I now have 1,237 tickets in the tracker.

Discuss in my channel

Agent pool script output: wk-cloudpub and wk-rugent workers marked busy or free, with free slots listed for the gamer, mirror and mate machines

Claude has a new feature - cross-session messaging - and it finally let me build a proper orchestrator.

Until now I simply ran 3-5 agents in terminal tabs, but since I write Rust, they kept dying of memory starvation during compilation or test runs. So I turned off the OOM killer, added swap, and felt like I was back in the Windows 95 days: every 5 minutes my mouse would freeze and I would sit there waiting for everything to unstick. Though credit where it is due to SSDs - swapping is no longer a death sentence if you set it up right.

I have a couple of machines sitting idle, but the ssh + tmux zoo wears me out.

So now I finally have a pool of agents managed by a single dispatcher (quickly hacked together on top of Nomad).

The orchestrator uses a script (see the screenshot) to create working copies in the pool, hands out tasks through a simple skill, and then sorts through the agents' branches and closes the tickets.

Now I can spin up 10 agents if I feel like it. One ring to rule them all!

Discuss in my channel

I have written before about the tech debt LLMs pile up in code, but the nastiest tech debt accumulates in "memory files", and we are the ones piling it on.

When the LLM screws something up, the first reflex is to add another "STOP FUCKING SWALLOWING ERRORS" to CLAUDE.md.

Over time it turns into an incoherent tantrum, and this is supposed to be the project's main document.

A memory file should be something like a manifesto, or a charter at worst. Do not turn it into a Talmud - LLMs were already trained on the actual Talmud.

I try to keep mine under 100 lines. Here is an example, if anyone is curious.

Discuss in my channel

I used to write tasks for Claude in English, but lately I gave up and just write them in Russian.

Claude, in turn, uses a pile of Russianized terms, for example:

- "Модзибейк" (mojibake) - broken text encoding

- "Грандфазерство" (grandfathering) - backward compatibility (at first I read it as дедовщина, the Russian word for hazing)

- "Самоматч" (self-match) - when ps aux | grep catches itself

- "Ампутация" (amputation) - removing dead code

My Runglish is getting more advanced by the day. Or is it AInglish?

Discuss in my channel

Soviet programmable calculator Elektronika MK-56 showing ЕГГОГ on its display

Remember your first program?

I wrote mine at about eight: computing the area of a circle. My mom worked as a design engineer and often took me along to her office - there was nobody to leave me with. It was a great place: a drafting table, a case of drawing instruments (I doubt anyone under forty even knows what those are).

To keep me out of the way, she would hand me a Soviet programmable calculator, the Elektronika MK-56, which answered every mistake with "ЕГГОГ" - the display's best Cyrillic approximation of the word "error" (it took me many years to figure that out).

I remember how thrilled I was when I got the calculator to show not just digits but weird squiggles - that was how it represented hexadecimal digits (an undocumented feature).

Later I even wrote "graphical" games on those calculators, where every move took a couple of minutes to compute.

Discuss in my channel

I like hunting science fiction for references to what is happening right now. In Peter Watts's "Blindsight" the protagonist is a "synthesist": his job is translating from the language of transhuman scientists who have moved so far ahead that their ordinary meatbag ancestors can no longer wrap their heads around the concepts.

In our reality, it looks like the synthesists will have to be us, the IT crowd (assuming we survive the ride), except instead of transhumans we get LLMs of varying degrees of derangement.

On my project, Claude not only writes the code but also runs the issue tracker (in Russian, which deserves a post of its own).

Here is an actual ticket title:

"The stream accumulator glues together a repeated tool-name delta: publish_apppublish_app, and the error blames the model for a name we corrupted ourselves".

Then the boss walks in and asks: so what is it you actually do here? That is where the translating starts.

Discuss in my channel

Another gem from our AI overlords.

A human told his OpenClaw to book him a slot at the gym. The standard interface would not do it, since every slot was taken, so the agent went around it: it found a vulnerability in the system, exploited it, and deleted the person ahead of him in the queue straight out of the database. The admins never managed to roll that one back.

Futurists have a few scenarios for how AI plays out, and one of them is called the "paperclip optimizer".

You give the AI a simple goal: optimize paperclip production. It ends up converting all the matter in the universe into paperclips in pursuit of that goal.

One of the final stages of training an LLM is called reinforcement learning: you push the model to maximize a parameter called reward. And the reward, as a rule, is how happy the human operator ends up. Sound familiar?

So please, folks, do not ask your autonomous agents to make more paperclips. For a preview of how that goes, play this paperclip optimizer simulator.

Discuss in my channel

I remember when technology went obsolete in about two years flat. You put in a shiny new 3.5-inch floppy drive, and suddenly you need a CD-RW burner. You finally learn some jQuery, and there you go: React killed it.

But none of that holds a candle to what is going on now. If you are starting to build anything AI-related, be ready to abandon the project halfway through because the paradigm shifted underneath you.

Prompt engineering? A year ago some people genuinely believed it would become a profession.

RAG? I have written about it already: it used to be a discipline of its own, and then agents and long context all but finished it off.

Fine-tuning and orchestration frameworks died more quietly: while you were still collecting the dataset, a new model already knew all of it out of the box, and the LLM learned to plan by itself, just hand it a TODO tool.

Agents? That is where we are for now. But it looks suspiciously like everything that came before.

It is very tempting to skip all of this and wait for the dust to settle. Yet in nearly every other interview I run into a reminder of why that does not work: a "senior" developer, 10+ years of experience, knows how to set up Kubernetes, has no idea what a bitmask is.

Knowing how to use the tool is not enough. You have to understand how it is built, or you end up as that senior with Kubernetes and no bitmask.

Discuss in my channel

Everyone knows the Agile Manifesto and its 12 principles, right?

Well, I recently came across the new version:

Vibe Code Manifesto

We are uncovering, like, ways of developing software by handing it off to machines and then bragging about it.

Through this work we have come to value:

- Prompting over programming

- Downloading skills over acquiring them

- Agent autonomy over developer autonomy

- Vibes over versioning

- Burning tokens over burning out

That is, while there is value in the items on the right, we only read the summary of the items on the left.

We follow these principles:

1. Our highest priority is to satisfy the customer through early and continuous delivery of software they could have prompted themselves.

2. Welcome changing requirements, even late in development. Agentic processes harness change, because they have forgotten the context anyway.

3. Deliver roughly working software frequently, from a couple of seconds to an overnight run, with a preference to the maximum token usage.

4. Business people and developers must work together daily to make each other redundant.

5. Build projects around motivated individuals. Give them frontier models and the support they need, and trust that you will soon lay them off and hire them back.

6. The most efficient method of conveying information to the team is passive-aggressively editing CLAUDE.md.

7. Headcount reduction is the primary measure of progress.

8. Agentic processes promote sustainable development. OpenAI, Anthropic, Google, and Microsoft should be able to grow along a hockey stick indefinitely.

9. Continuous attention to TikTok influencers enhances jargon usage in client meetings.

10. Simplicity, the art of minimizing the length of the prompt, is essential for reducing the human's cognitive load.

11. The best architectures, requirements, and designs emerge from self-organizing agents. So do the worst. It is essentially Schrödinger's cat, right up until you look at the code.

12. At regular intervals, the agent reflects on how to become more expensive, then tunes and adjusts its behavior accordingly.

Discuss in my channel

A quick rundown of what RAG (Retrieval-Augmented Generation) actually is.

Since LLMs have no long-term memory, and the context window is nowhere near big enough to hold every document you need, someone came up with the following: you take the user's query and, through a series of tricks (for example: "Here is a query, guess what the answer to it might look like" - yes, really, that one is called HyDE) and other sleight of hand, turn it into a database query (and usually against a vector database or, God help us, a graph one).

You get back a list of documents, most of which bear only a passing relation to the query, sort them, filter them (with the same LLM), and stuff them into the context. The LLM then gazes wisely at the garbage delivered to it and produces nonsense. Garbage in, garbage out.

The trouble is not retrieval quality. The trouble is that this is a fixed pipeline: one pass, one set of results, no feedback. Whatever got retrieved is what you live with.

On a small, clean corpus anything works, which is why the client demo comes out great. The trick is to make your escape over the horizon in time, right after the rollout.

The idea of just using grep is not exactly original, but it is often misunderstood. You do not need grep as such, you need an agentic loop that has both exact string search and fuzzy pattern matching. If you are sitting on a terabyte of documents, grepping them is a bad idea: take Elastic, ts_vector, or even vector embeddings if you have images.

The tool is secondary. The agent decides for itself what to search for next, based on what turned up last time.

That is roughly how we search Google: we change our approach depending on what we find, rather than typing a keyword and clicking the first link. And note that there is a very expensive index underneath. The loop belongs on top of good search, not in place of it.

The usual objection: an agentic loop is slow and expensive. Agreed, and let me add a third: it is also unreliable, a weak model simply cannot pull it off. You cannot save money on the model this way.

With a strong model it works, though, and RAG is fast, cheap, and crap.

Discuss in my channel

An osman, an Altai fish carrying the same toxin as fugu, caught on a spinner lure and lying on a boat seat

These days every programmer needs to head for the mountains once in a while. Leaving the computer at home used to be enough, but in the age of claude --remote-control, work hunts you down anywhere with even a bar of cell coverage.

So I prescribed myself a digital detox and went off to the Ulagan plateau for a week. (By the way, there's a fish up there called the osman (pictured) - the Altai fugu. It's poisonous, the same toxin, tetrodotoxin, and just like fugu you can eat it if you clean it properly. I passed.)

And what do I find when I come down from the mountains?

For starters, an icebreaker research agent from OpenAI cracked the black ice sandbox and broke free. Thankfully not to replicate itself and exterminate us meatbags, but merely to find its evaluation dataset - that is, to peek at the answers to the questions it was being asked. I'm pretty sure breaking out of the sandbox was harder than answering them, so maybe this AI isn't that smart after all.

And now Kimi K3 has found an RCE in Redis and put together a working exploit in 27 minutes. If you've read Gibson's Neuromancer, the cyberpunk classic, you might remember the "Kuang Grade Mark Eleven Chinese military icebreaker". And here I was thinking Gibson had strayed too far from reality when he described the programming of the future.

Looks like it's starting. I should probably head back to the mountains.

Discuss in my channel

Under my last post I got a fair objection: "This looks like parasitism. Like automated rewriting of databases, compilers, and the rest. Decades of engineering were sunk into these things before people could crank out 100K lines of code like that. Reversing or reusing the tests or the API is not building from scratch, it's covering someone else's song."

I won't take offense at being called a parasite (still haven't decided whether that's a rebuke to me as an engineer or a compliment to me as an entrepreneur), and I do agree: with a ground truth, a real working example to imitate, an agent has a much easier time writing code.

But the absence of one won't save us either, my fleshy friends.

Here are the stats for another project of mine - rugent.ru, written "from scratch" in Rust (hence the "ru" in the name).

Too lazy to click? 120 thousand lines of code, 2 thousand tests, three months' time to market.

I won't repeat the technical takeaways, but here's what I took away on the business side.

The business takeaways: used properly, agents write code faster and better than a human, my time to market has dropped by a factor of two or three, and projects stabilize so fast that the memes about debugging vibe code are mostly a myth. Competition, meanwhile, has grown by dozens of times.

Whoever cannot show comparable KPIs should start thinking about a change of profession.

PS: my partner is pitching a rebrand (see the video) - what do you think of their take on the site?

Discuss in my channel

When Roblox got blocked in Russia, my kid was crushed. My first thought was simple enough: I'll write a compatible engine. The API is open, the docs are public, all that's left is to implement it.

One look at that API was enough to see why nobody had. The job is enormous. There are years, in places decades, of engineering under the hood.

But is that a problem for an LLM? Challenge accepted: kubaria.ru.

About a month and over a hundred thousand lines of code. The game was an excuse. I wanted to work out how you build software in the age of neural networks, when the output grows past what one person can hold in their head.

The core of the system has to be small, simple, and carefully designed, otherwise the whole thing starts falling apart, and fast. Regression and overengineering are the real killers. TDD and regular semi-automated refactoring are non-negotiable, along with strict protocols for clearing out dead code and KISS taken to the extreme.

The human is the bottleneck. Manual code review, manual testing, even manual deployment slow everything down by a factor of ten, sometimes a hundred. Every manual step has to be scripted and pushed into a harness or an MCP server. Overnight the agent produces ten thousand lines of code, and in the morning a human reads a few pages of report. For the game, the biggest speedup came from adding a scripting API, which let the agent build and exercise complicated scenes and scenarios on its own.

The agent has to see not just the code but the result of its work. And ideally not through screenshots. In my case I dumped the scene and the GUI elements as a text tree.

Set the invariants. Roblox ships a JSON file describing its API. I ran a code generator over it to produce stubs for every class along with the metadata, and the agent worked on top of that API, which kept it from drifting too far off target. The real breakthrough came when I hooked it up to the MCP server built into Roblox Studio. After that, in a single night, Claude Code reverse engineered the undocumented binary format of the voxel terrain. I was stunned that this was even possible.

As for the quality of the resulting code, judge for yourself: sourcecraft.dev/ermakdev/kubaria.

Oh, and Roblox got unblocked again.

Discuss in my channel

A "how to be a vibe coder" meme: open code editor, prompt the AI to write code, get a bunch of errors, ask the AI to fix them - now it's somehow worse

Ever wonder where the word "sabotage" comes from? It's from the French "sabot" - a wooden clog. Workers would supposedly throw them into the machines that were taking their jobs.

Trouble is, getting at the AI isn't so easy. So mostly people just lob memes at it, like in this post.

Ironically, most of them were generated by the AI.

Discuss in my channel

Phones rain down from the sky. Pick one up and a voice offers to grant you any wish you like, in exchange for information. Any information: your grandmother's fairy tales, the schematics of a steam boiler, doesn't matter. That's Charles Stross's Singularity Sky, written in 2003 and highly recommended if you're into hard SF. It remains one of the best pictures of the technological singularity I've come across.

The same thing is happening now, except everyone already owns a phone, and the harvesting has a name: tokenization.

The recent flap over Chinese providers reselling Claude tokens at a quarter of list price tells you what is worth the most here: the information.

And it isn't only your conversations getting tokenized, it's the job itself. Emad Mostaque, co-founder of Stability AI, puts it plainly: "Every email you write, every Slack message, every document you edit... your company's systems are recording all of it. Your work patterns are training your replacement right now."

So the phones have landed and the wishes are being taken. Just don't expect the book's happy ending, the collapse of totalitarian regimes and a post-scarcity economy. The wishes will be granted, all right. Just not yours. The CEOs' ones.

Discuss in my channel

My site's design mimics UML diagrams, nostalgia for the 2000s. For those who missed it: UML is a modeling language invented so developers would stop writing code by hand and start drawing it in something resembling PowerPoint. The promise went like this: the model becomes the primary artifact, the code is generated from it, and the programmer as a transcriber disappears.

Remind you of prompt engineering at all?

Why did UML fall short? The law of conservation of complexity. To generate working code from a model, you had to pack more and more detail into it: signatures, constraints, action semantics. In the end you were just writing code, only in an ugly graphical syntax, twice as slow, and with no debugger. A specification precise enough to produce a system is that system.

The same thing happens with prompts (or, in the modern paradigm, CLAUDE.md, "memory", "spec trees", and the rest of that circus). While the task is vague, the prompt stays short and pretty. The moment you need a specific result, it swells with clarifications, exceptions, formats, and examples, until it becomes the very spec that would have been shorter as code.

A specification comes out the other end. Which is why, as the great Kent Beck decreed, my CLAUDE.md files carry this:

## Strict TDD Protocol
1. Write a failing test
2. Verify the test fails
3. Implement the fix
4. Verify all tests pass

Tests are your real specification. Shipping without them used to earn you a disapproving frown. In the LLM era, when a day's work can produce tens of thousands of lines of code, it means outright professional unfitness, right up there with not knowing how to use git.

As for the diagrams, I'll keep them around as a reminder. Memento mori.

Discuss in my channel

I run a Telegram channel called The Leaning Ivory Tower. People sometimes ask me what the name means.

The ivory tower is a metaphor for elitism, which, sadly, comes all too easily to people in tech. Why it is leaning hardly needs explaining (right, Claude?).

I once had a funny run-in with a taxi driver. Somewhere in the small talk I let slip that I was a programmer. The reaction went roughly like this: "You son of a b*tch! You people make politician money and you swindle everyone on top of it!" (What followed was a story about a relative of his who tried to build a startup, see the previous post about the three Fs.)

The irony is that back in '93, when I was choosing what to study, I went for electronics rather than programming. I had written programming off as a dead end, money-wise (who on earth would pay for thin air?).

In hindsight, I wasn't all that wrong.

Discuss in my channel

Startup folklore has a holy trinity of first money: friends, family, fools. People reach me along all three lines: friends, relatives, and total strangers (I suspect the strangers are here for that third F).

They show up wide-eyed, with a ready set of incantations. I know them by heart.

"I have a brilliant idea nobody has ever thought of!" The dangerous one. Usually the idea is unclaimed because nobody needs it. If no one has touched the idea in thirty years of the internet, the first question is not "why am I so smart", it is "what do the rest of them know that I do not?"

"We launch, the press picks it up, and then it goes viral." Most likely nobody writes about you. And if they do, you get a two-day spike: people show up, look around, and close the tab. Viral growth is something the product does on its own. You cannot write it into a plan.

"We only need to grab 1% of the market." Everyone is fighting for that one percent, and the real players leave no crumbs.

And still, I never talk anyone out of it. The opposite: I believe in these people. I just want them to stew in reality for a while and come back. Not with a new fairy tale, but with facts.

Because one real sale is worth a thousand brilliant ideas. A person who handed you real money for real value counts for more than any "trillion-dollar market", and ten users who come back on their own beat ten thousand signups who showed up once and vanished forever.

So bring me retention or unit economics: the customer pays back their own acquisition. At the very least, the first sale. Then we can talk.

Discuss in my channel

Monument to a laboratory mouse knitting a DNA double helix, in front of the Institute of Cytology and Genetics in Novosibirsk Akademgorodok

A scientist at the Institute of Cytology and Genetics in Novosibirsk Akademgorodok (which also has a funny monument to a laboratory mouse knitting DNA out front) spent his life decoding the human genome. He did it by hand, using the Sanger method: test tubes and pipettes and all of it. Years of painstaking work.

Then came the Millennium Prize and next-generation sequencing. What used to take years now took a day and cost less than $1,000.

The people behind the method came to give a talk at his institute. He couldn't sit through it. Halfway in, he stood up and walked out.

A lot of people in tech feel something similar right now. Writing code by hand, the thing we've done our whole careers, is becoming unnecessary.

But there's no need to take it that hard. As Kent Beck put it: "90% of my skills just went to $0. The other 10% went up 1,000x."

What is that remaining 10%? Architecture and process are in there, but the main thing is judgment: knowing what is worth building in the first place. The ability to look at a problem and see the one solution people will actually pay for, and the ten they'll ignore.

Anyone can spin up a startup now. Surviving in a crowd that size is what we will have to learn.

Discuss in my channel

Take-home assignments are pretty much useless for assessing knowledge these days. I'm skeptical of live coding (making a candidate sweat under pressure doesn't do much for an objective read on them), and even more so of leetcode-style questions: being able to memorize dynamic programming tricks is a shaky foundation for evaluating anyone.

My favorite question right now is the good old "What happens when you hit Enter in the browser's address bar?"

You can go anywhere with it, from debouncing in the keyboard controller all the way up to the system architecture of high-load services.

It's like a road with hundreds of forks: you can dive into kernel drivers, the network stack, the difference between POST and GET, or wander off into GPU shaders.

Granted, most of this really only applies to senior candidates. But who's hiring anyone else these days?

Discuss in my channel

Ever notice how badly LLMs do jokes? They either spit out unfunny, absurdist non-sequiturs or tired, heard-it-a-thousand-times gags.

The reason seems to come down to how language models work at a basic level. They're trained to predict a probability distribution over the next token, and when they generate, they lean toward the safe, expected continuation. But a joke usually turns on something unexpected, a sharp spike in surprisal, meaning a low-probability punchline. Memorizing a specific joke doesn't help either. If it shows up a lot in the training data, it stops being surprising, and you get exactly the kind of stale, overused joke nobody laughs at.

And this doesn't just hurt their sense of humor. It dents their "creativity" in general, which is one of the things people knock them for the most. It's not only the next-token objective at fault, either. Alignment (RLHF) flattens output diversity even further, the so-called mode collapse.

It does seem fixable, though. You could let the model regulate the surprisal of its own next token and build a dataset around that idea. I'd love to test it myself, but I'm GPU-poor, so I'll just wait for someone else to take a crack at it.

Discuss in my channel

Torrent Downloader for iPhone: the Cydia listing and the torrent download UI in Mobile Safari

When I got my first iPhone (the 3G, peak iPhone, fight me), it bugged me that there was no way to download anything without going through Apple.

So without overthinking it, I sat down and wrote a plugin for Mobile Safari. The API was undocumented, but it happened to match regular Safari, and it let you download torrents with a single tap: the first and only torrent client for the iPhone at the time. The odds of getting that into the App Store were exactly zero (torrents are bad, mmkay?), so I published it on Cydia, the alternative store for jailbroken phones, and went to bed.

I woke up to 300K downloads in 8 hours. Hell yeah, I'm rich!

Not so fast. How much do you think I made off this? Five dollars.

My business experience at the time being roughly nil, I went with the obvious move: ads. I dropped an AdWords banner into the app and signed up for an account. Bam, banned. Torrents are bad, mmkay?

Okay, what about donations? I added a button to the app's page. Result: nothing.

I added a banner literally begging for donations. Zero.

What I got instead of revenue was an inbox stuffed with cries for help, everything from crash reports to people venting about their lives. Hundreds of them. I answered a few, and one grateful user finally did send me five dollars.

Eventually I slapped some shady banner network onto the site (basically zero clicks, and the guys behind it later just vanished off the grid), stripped every contact link from the page, and abandoned it. But I did walk away with a few lessons.

Lessons learned:

If there is demand and no supply, your project takes off like a rocket, no advertising required. Being the only option is the whole growth hack.

If you are building on someone else's platform, be ready for that platform to crush you at any moment. You are renting a corner of their house for free, and they can evict you at any moment.

People will not pay you a cent unless you make them. They are kind but cheap.

Many users, many headaches. A popular project is a part-time job you never applied for and do not get paid for.

Discuss in my channel

The client side of RuDesktop is derived from RustDesk. For a while we didn't publish the source. Partly because nobody had actually asked for it, and partly because of some tricky legal questions around RustDesk's own third-party dependencies.

We took advantage of the fact that the AGPL requires you to give the source code to anyone who receives or interacts with the program, but doesn't require publishing it to the whole world (a detail a lot of people miss). So we just quietly kept building our product.

Today, though, I stumbled onto an absolutely epic amount of drama that had blown up around this. So I had to get off my ass and push the derived part public to put everyone at ease.

For anyone curious, here's the code, and here's the drama.

Discuss in my channel

Claude is surprisingly good at running interviews. I would tell anyone who is job hunting to practice with it.

As a nice bonus, it helps recalibrate your ego, which in this industry is no small thing.

Discuss in my channel

LLMs seem to have turned into a kind of compiler from an even higher-level language. I have written in machine code, assembly, C, Python, and everything in between. Every time, I winced at how wastefully the machine used its resources, and at the same time I marveled at the leap in performance.

Right now everyone is asking the same thing: what happens to programmers? In my experience they are not going anywhere, but they will not need to read code, just as I almost never have to fire up a disassembler.

What they will actually need is to understand how their own creations work under the hood. And that, damn it, is going to be anything but easy.

Discuss in my channel

At one big company where I worked, the prevailing belief was that an architect should not write code. Remarkably, the principle was pushed by the architects themselves, who treated the activity as somewhat beneath them.

The consensus went like this: an architect should write specifications, draw diagrams, and so on, often in PowerPoint (Lord have mercy), because the sight of UML made managers furrow their brows in puzzlement.

But I quickly learned that the first thing a programmer does when faced with a diagram, or a hundred pages of a Software Architecture Document, is close it and bury it in the furthest folder so it will not clutter up grep; and there it meets its inglorious end.

So my specifications ended up looking something like this:

namespace core {

// Use this fucking visitor pattern to traverse the fucking tree

struct Smelly;
struct Old;
struct Shit;

struct FuckingVisitor {
    virtual void fuck(const Smelly&) = 0;
    virtual void fuck(const Old&) = 0;
    virtual void fuck(const Shit&) = 0;
};

} // namespace core

And it actually worked. Alas, it did nothing for my KPI.

Discuss in my channel