<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xmlns:xhtml="http://www.w3.org/1999/xhtml">
  <title>Blog — Stefan Ostermann</title>
  <subtitle>Blog posts by Stefan Ostermann on software development, AI infrastructure, and scalable systems.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://www.stefanostermann.de/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://www.stefanostermann.de/blog.html"/>
  <id>https://www.stefanostermann.de/feed.xml</id>
  <updated>2026-08-23T00:00:00Z</updated>
  <rights>© 2026 Stefan Ostermann</rights>
  <author>
    <name>Stefan Ostermann</name>
    <email>me@stefanostermann.de</email>
  </author>
  <entry xml:lang="de">
    <title>Ein Prompt für den Agenten über Nacht und ein Strategiespiel</title>
    <link rel="alternate" href="https://www.stefanostermann.de/de/one-prompt-one-night-one-strategy-game.html"/>
    <id>https://www.stefanostermann.de/de/one-prompt-one-night-one-strategy-game.html</id>
    <updated>2026-08-23T00:00:00Z</updated>
    <published>2026-08-23T00:00:00Z</published>
    <summary>Zwei Experimente mit einem lokalen Modell mit viel Freiheit über Nacht.</summary>
    <content type="html">&lt;p&gt;Ich war unterwegs, und das einzige Gerät im Gepäck war ein MacBook Air M4 mit 24 GB Arbeitsspeicher. Alibaba hatte Qwen 3.8 27B am Tag vorher veröffentlicht. Ich war den ganzen Tag hibbelig, abends dann endlich etwas Freizeit. Also habe ich eine niedrige Quantisierung geladen (3 Bit) und lande bei etwa zwei Tokens pro Sekunde. Unbrauchbar für Chat / Live Vibecoding, aber gut genug für einen Test im pi-Harness:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;„Can you create a html / js / css based space based strategy game?&amp;quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Ich hab&apos;s dann einfach laufen lassen.&lt;/p&gt;
&lt;p&gt;Es lief die ganze Nacht und einen großen Teil des nächsten Tages. Herausgekommen ist ein spielbares rundenbasiertes Strategiespiel: eine Sternenkarte mit eigenen und fremden Systemen, Mineralien, Energie und Bevölkerung als Wirtschaft, vier Schiffsklassen mit unterschiedlichen Kosten und Reichweiten, ein Forschungsbaum. Es sah aus wie etwas, das ich als Teenager gerne gespielt hätte.&lt;/p&gt;
&lt;h3&gt;Zum Nachdenken&lt;/h3&gt;
&lt;p&gt;Das Modell hat nicht nach dem Spiel aufgehört. Es hat eine Simulation geschrieben, um das Spiel durchzuspielen (eine Seite von einer KI gesteuert, die andere so, wie ein Mensch spielen würde) und danach ausgerechnet, wie leicht es zu schlagen ist.&lt;/p&gt;
&lt;p&gt;Es hat nicht gewartet, bis mir die Balance-Probleme auffallen. Es hat sie selbst gefunden, nachts um drei auf einem Notebook.&lt;/p&gt;
&lt;p&gt;Das zweite Experiment, wieder zu Hause mit der Keller-KI, wo dasselbe Modell viel komfortabler läuft: Ich habe ihm einen einfachen Planetengenerator gegeben, den ich vor sieben Jahren geschrieben hatte. Wieder ein einfacher Prompt. Zurück kam ein Strategiespiel mit prozedural generiertem Universum, erzählerischen Momentaufnahmen zwischen den Zügen, eigenständigen Charakteren und Handels- und RPG-Mechaniken in einem Game Loop. Zweimal habe ich nachprompten müssen, um kleine Fehler zu glätten. Entstanden ist &lt;a href=&quot;https://thoster.net/ts-universe/&quot;&gt;TS UNIVERSE — The Long Drift&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Ein Sternensystem aus TS UNIVERSE — The Long Drift: Zeov-84, ein Roter Zwerg mit einer Relaisstation, im Territorium The Mycelia.&quot; src=&quot;../img/ts-universe.8af1405b.webp&quot; width=&quot;1600&quot; height=&quot;939&quot; srcset=&quot;../img/ts-universe-800.77e56f63.webp 800w, ../img/ts-universe.8af1405b.webp 1600w&quot; sizes=&quot;(max-width: 767px) calc(100vw - 3rem), 46rem&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h3&gt;Der Haken&lt;/h3&gt;
&lt;p&gt;Bei kleiner Hardware: Thinking bei Qwen 3.8 27b dauert ewig. Ein Ein-Prompt-Auftrag antwortet in Stunden, selbst wenn das Prompt und die Aufgabe klein sind. Für alles Interaktive taugt das Modell mit zu schwacher Hardware nicht: Ich sitze nicht davor und schaue beim Nachdenken zu, während ich überlege, was ich als Nächstes tippe.&lt;/p&gt;
&lt;p&gt;Für Batchverarbeitung ist es damit trotzdem sehr brauchbar.&lt;/p&gt;
&lt;h3&gt;Erkenntnisse&lt;/h3&gt;
&lt;p&gt;Wir bewerten Modelle nach Benchmarks und Chat. Das belohnt schnelle und gefällige Antworten. Gemessen habe ich aber hier, wie viel in sich geschlossene, prüfbare Arbeit ein Modell allein tragen kann, bevor es mich braucht. Die Antwort wäre vor wenigen Monaten „kaum etwas&amp;quot; gewesen, und heute ist sie „ein komplettes Spiel&amp;quot;.&lt;/p&gt;
&lt;p&gt;Im Batch über Nacht kann es auch auf eigentlich zu schwacher Hardware brauchbare Ergebnisse liefern. Im Chat braucht man schon etwas stärkeres. Aber nicht nur das Modell spielt eine Rolle, auch der Harness, hier war das Pi. Er hat dafür gesorgt, dass im agentischen Loop ein besseres Ergebnis herausgekommen ist als das mit klassischen Entwicklungsumgebungen + KI Chat der Fall gewesen wäre.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.stefanostermann.de/one-prompt-one-night-one-strategy-game.html"/>
  </entry>
  <entry xml:lang="en">
    <title>A Prompt for the Agent Overnight, and a Strategy Game</title>
    <link rel="alternate" href="https://www.stefanostermann.de/one-prompt-one-night-one-strategy-game.html"/>
    <id>https://www.stefanostermann.de/one-prompt-one-night-one-strategy-game.html</id>
    <updated>2026-08-23T00:00:00Z</updated>
    <published>2026-08-23T00:00:00Z</published>
    <summary>Two experiments with a local model and a lot of freedom overnight.</summary>
    <content type="html">&lt;p&gt;I was travelling, and the only machine with me was a MacBook Air M4 with 24 GB of memory. Alibaba had released Qwen 3.8 27B the day before. I was on edge about it all day; in the evening, finally, some free time. So I downloaded a low quantisation (3 bit) and landed at around two tokens per second. Useless for chat or live vibe coding, but good enough for one attempt in the pi harness:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;quot;Can you create a html / js / css based space based strategy game?&amp;quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I just left it running.&lt;/p&gt;
&lt;p&gt;It went all night and a good part of the next day. What came out was a playable turn-based sector strategy game: a star map with systems you own and systems you do not, minerals, energy and population as the economy, four ship classes with different costs and ranges, a research tree. It looked like something I would have loved to play as a teenager.&lt;/p&gt;
&lt;h3&gt;The part I keep thinking about&lt;/h3&gt;
&lt;p&gt;The model did not stop at building the game. It wrote a simulation to play the game to the end (one side controlled by an AI, the other playing the way a human would) and then worked out how easy it is to beat.&lt;/p&gt;
&lt;p&gt;It did not wait for me to discover the balance problems. It found them, on its own, at three in the morning, on a laptop.&lt;/p&gt;
&lt;p&gt;A second experiment, back home on the basement machine, where the same model runs much more comfortably: I handed it a simple planet generator I wrote seven years ago. Again one simple prompt. What came back was a strategy game with a procedurally generated universe, narrative snapshots between turns, distinct characters, and trading and RPG mechanics in the game loop. I prompted twice more to iron out small issues. The result is &lt;a href=&quot;https://thoster.net/ts-universe/&quot;&gt;TS UNIVERSE — The Long Drift&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;A star system from TS UNIVERSE — The Long Drift: Zeov-84, a red dwarf with a relay station, in the territory The Mycelia.&quot; src=&quot;img/ts-universe.8af1405b.webp&quot; width=&quot;1600&quot; height=&quot;939&quot; srcset=&quot;img/ts-universe-800.77e56f63.webp 800w, img/ts-universe.8af1405b.webp 1600w&quot; sizes=&quot;(max-width: 767px) calc(100vw - 3rem), 46rem&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h3&gt;The catch&lt;/h3&gt;
&lt;p&gt;On small hardware: thinking in Qwen 3.8 27B takes forever. A one-shot prompt answers in hours, even when the prompt and the task are small. That makes the model on too weak a machine useless for anything interactive: I will not sit and watch it deliberate while I decide what to type next.&lt;/p&gt;
&lt;p&gt;For batch processing it is nevertheless very useful.&lt;/p&gt;
&lt;h3&gt;Takeaways&lt;/h3&gt;
&lt;p&gt;We judge models on benchmarks and on chat, which rewards fast and agreeable answers. But what I measured was how much self-contained, testable work a model can carry on its own before it needs me. The answer a few months ago would have been &amp;quot;very little&amp;quot;, and today it is &amp;quot;a whole game&amp;quot;.&lt;/p&gt;
&lt;p&gt;In batch, overnight, it can deliver useful results even on hardware that is arguably too weak. In chat you need something stronger. But it is not just the model — the harness plays a part, too; here it was Pi. In the agentic loop it produced a better result than classic development environments plus AI chat would have.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="de" href="https://www.stefanostermann.de/de/one-prompt-one-night-one-strategy-game.html"/>
  </entry>
  <entry xml:lang="de">
    <title>Zwei Grafikkarten im Keller</title>
    <link rel="alternate" href="https://www.stefanostermann.de/de/two-gpus-in-the-basement.html"/>
    <id>https://www.stefanostermann.de/de/two-gpus-in-the-basement.html</id>
    <updated>2026-08-19T00:00:00Z</updated>
    <published>2026-08-19T00:00:00Z</published>
    <summary>Warum aus einem Feierabendprojekt ein Grund geworden ist, meinen Job aufzugeben — und was ich ab Oktober mache.</summary>
    <content type="html">&lt;p&gt;In meinem Keller steht ein Rechner mit zwei Grafikkarten. Darauf arbeiten Tag und Nacht KI-Modelle. Niemand kann sie limitieren, niemand die Preise erhöhen, niemand mitlesen.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Eine der beiden Grafikkarten: eine NVIDIA GeForce RTX 5090 Founders Edition, frisch ausgepackt.&quot; src=&quot;../img/rtx-5090-founders-edition.354e3c17.webp&quot; width=&quot;1600&quot; height=&quot;1205&quot; srcset=&quot;../img/rtx-5090-founders-edition-800.fd21f9ce.webp 800w, ../img/rtx-5090-founders-edition.354e3c17.webp 1600w&quot; sizes=&quot;(max-width: 767px) calc(100vw - 3rem), 46rem&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;p&gt;Angefangen hat das als Feierabendprojekt. Ein bisschen Hardware, ein paar offene Modelle, viel Ausprobieren. Inzwischen ist es einer von mehreren Gründen, warum ich meinen Job aufgegeben habe.&lt;/p&gt;
&lt;h3&gt;Ab dem 01.10.2026 bin ich selbstständig&lt;/h3&gt;
&lt;p&gt;Nach vielen Jahren als Führungskraft in der Software- und Produktentwicklung, zuletzt als Head of Development, berate ich Unternehmen künftig zu privater und agentischer KI: welche Hardware, welche Modelle, welche Architektur. Damit KI wirklich produktiv wird und die Daten dabei im Unternehmen bleiben.&lt;/p&gt;
&lt;h3&gt;Warum jetzt?&lt;/h3&gt;
&lt;p&gt;Weil aus dem Nebenprojekt eine Überzeugung geworden ist: Der Weg zu KI führt nicht zwangsläufig über einen US-Cloud-Anbieter.&lt;/p&gt;
&lt;p&gt;Vieles, wofür heute pro Nutzer und Monat gezahlt wird, läuft lokal. Schnell genug, günstig genug und nachvollziehbar. Nicht alles, für manche Aufgaben sind die großen Frontier-Modelle weiterhin das richtige Werkzeug. Aber der Anteil dessen, was auf eigener Hardware sinnvoll läuft, ist im letzten Jahr deutlich größer geworden, und die Rechnung geht für viele mittelständische Unternehmen inzwischen auf.&lt;/p&gt;
&lt;h3&gt;Näher am Metall&lt;/h3&gt;
&lt;p&gt;Ich werde dabei wieder deutlich operativer arbeiten. Möglich macht das erst die Technik selbst: KI-Agenten übernehmen im Hintergrund die Fleißarbeit. Das verschafft mir Zeit für den Teil, der jetzt der wichtigste ist: die Arbeit mit Menschen.&lt;/p&gt;
&lt;p&gt;Ich starte allein. Mit dem klaren Plan, gemeinsam mit einem Partner zu wachsen.&lt;/p&gt;
&lt;h3&gt;Ab Oktober habe ich Kapazität&lt;/h3&gt;
&lt;p&gt;Für Projekte, Sparring und Kooperationen. Und wenn Ihnen beim Lesen jemand einfällt, der genau vor diesen Fragen steht: Ich freue mich über eine Vorstellung.&lt;/p&gt;
&lt;h3&gt;Wie das für ein Unternehmen aussieht&lt;/h3&gt;
&lt;p&gt;Wie ein solches Projekt abläuft, vom Audit über die Hardware und die Modelle bis zur Übergabe, habe ich aufgeschrieben: &lt;a href=&quot;private-ai.html&quot;&gt;KI auf eigener Hardware&lt;/a&gt;.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.stefanostermann.de/two-gpus-in-the-basement.html"/>
  </entry>
  <entry xml:lang="en">
    <title>Two Graphics Cards in the Basement</title>
    <link rel="alternate" href="https://www.stefanostermann.de/two-gpus-in-the-basement.html"/>
    <id>https://www.stefanostermann.de/two-gpus-in-the-basement.html</id>
    <updated>2026-08-19T00:00:00Z</updated>
    <published>2026-08-19T00:00:00Z</published>
    <summary>How an after-hours project became one reason I left my job — and what I&apos;m doing from October.</summary>
    <content type="html">&lt;p&gt;There is a machine in my basement with two graphics cards. AI models work on it day and night. Nobody can rate-limit them, nobody can raise the price, nobody reads along.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;One of the two graphics cards: an NVIDIA GeForce RTX 5090 Founders Edition, fresh out of the box.&quot; src=&quot;img/rtx-5090-founders-edition.354e3c17.webp&quot; width=&quot;1600&quot; height=&quot;1205&quot; srcset=&quot;img/rtx-5090-founders-edition-800.fd21f9ce.webp 800w, img/rtx-5090-founders-edition.354e3c17.webp 1600w&quot; sizes=&quot;(max-width: 767px) calc(100vw - 3rem), 46rem&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;p&gt;It started as an after-hours project. Some hardware, a few open models, a lot of trial and error. It is now one of several reasons I left my job.&lt;/p&gt;
&lt;h3&gt;From 1 October 2026 I am self-employed&lt;/h3&gt;
&lt;p&gt;After many years leading software and product development, most recently as Head of Development, I will be advising companies on private and agentic AI: which hardware, which models, which architecture. So that AI actually becomes productive and the data stays inside the company.&lt;/p&gt;
&lt;h3&gt;Why now?&lt;/h3&gt;
&lt;p&gt;Because a side project turned into a conviction: the road to AI does not have to run through a US cloud provider.&lt;/p&gt;
&lt;p&gt;A lot of what is billed per user per month today runs locally. Fast enough, cheap enough, and auditable. Not everything, for some tasks the large frontier models are still the right tool. But the share of work that runs well on your own hardware has grown considerably over the past year, and for many mid-sized companies the maths now works out.&lt;/p&gt;
&lt;h3&gt;Closer to the metal&lt;/h3&gt;
&lt;p&gt;I will be working far more hands-on again. What makes that possible is the technology itself: AI agents take over the legwork in the background. That frees up time for the part that matters most now — working with people.&lt;/p&gt;
&lt;p&gt;I am starting solo, with a clear plan to grow together with a partner.&lt;/p&gt;
&lt;h3&gt;From October I have capacity&lt;/h3&gt;
&lt;p&gt;For projects, sparring, and partnerships. And if someone came to mind while you were reading this, someone facing exactly these questions: I would be glad for an introduction.&lt;/p&gt;
&lt;h3&gt;What that looks like for a company&lt;/h3&gt;
&lt;p&gt;I have written down how such a project runs, from the audit through the hardware and the models to the handover: &lt;a href=&quot;private-ai.html&quot;&gt;AI on your own hardware&lt;/a&gt;.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="de" href="https://www.stefanostermann.de/de/two-gpus-in-the-basement.html"/>
  </entry>
  <entry xml:lang="de">
    <title>Ein Agent arbeitet, während ich schlafe</title>
    <link rel="alternate" href="https://www.stefanostermann.de/de/an-agent-while-i-sleep.html"/>
    <id>https://www.stefanostermann.de/de/an-agent-while-i-sleep.html</id>
    <updated>2026-08-17T00:00:00Z</updated>
    <published>2026-08-17T00:00:00Z</published>
    <summary>Der unterschätzte Vorteil von lokaler KI ist nicht Tempo oder Qualität. Es sind die acht Stunden, in denen niemand hinschauen muss.</summary>
    <content type="html">&lt;p&gt;Öffentliche KI hat einen Gebührenzähler. Jede Anfrage summiert sich, ein langer Agentenlauf am Abend knabbert an dem Kontingent, das ich morgen für die eigentliche Arbeit brauche. Lokale KI nicht.&lt;/p&gt;
&lt;p&gt;Diesen Unterschied habe ich neulich bewusst genutzt. Vor dem Schlafengehen habe ich meinem Inferenz-Server den Job gegeben, den ich abends nicht mehr schaffen würde: die fehlenden CRUD-Endpoints meiner aktuellen App, inklusive der Tests, die hätten dabei sein müssen. Ein Feld löschen soll es auch wirklich auf null setzen, eine Organisation muss samt ihrer Standorte und Termine verschwinden, und nichts davon war abgedeckt. Dann bin ich schlafen gegangen.&lt;/p&gt;
&lt;h3&gt;Nach dem Aufwachen&lt;/h3&gt;
&lt;p&gt;Ein Branch mit Commits, die Endpunkte auf dem Server implementiert, die Tests angehängt in jede betroffene Testdatei. Eine Aufgabenliste, die bei eins von fünf anfing und leer endete. Auf llama.cpp mit Qwen 3.8 27B.&lt;/p&gt;
&lt;p&gt;Spannend finde ich den Preis. Die Kiste zieht unter Last rund 650 Watt. Acht Stunden davon sind etwa fünf Kilowattstunden, das sind bei meinem Tarif circa 1,50 Euro für eine Nacht Arbeit. Kein Tokenzähler, kein Rate-Limit, kein „wöchentliches Limit erreicht&amp;quot; mitten im Refactoring. Ich kann etwas starten, das in einer Stunde fertig wird, und etwas, das drei Tage läuft.&lt;/p&gt;
&lt;h3&gt;Nachteile&lt;/h3&gt;
&lt;p&gt;Ein fertiger Branch heißt nicht: fertig. Ich lese den Diff morgens, denn ein Agent, der acht Stunden ungestört war, hatte acht Stunden Zeit, überzeugend falsch zu liegen. Langsames Review ist sozusagen der Strafzoll unbeaufsichtigter Arbeit.&lt;/p&gt;
&lt;p&gt;Vorgestern bin ich mit einer Lösung für ein Problem aufgewacht, das ich schlecht beschrieben hatte. Das Modell war äußerst gründlich in der falschen Richtung. Das hat mich eine Nacht gekostet.&lt;/p&gt;
&lt;p&gt;Lokal auf meinem Notebook ist das Modell unerträglich langsam, wenn ich in Echtzeit warte. Wenn ich aber einfach ins Bett gehe, dann ist&apos;s erträglich!&lt;/p&gt;
&lt;h3&gt;Offene Fragen&lt;/h3&gt;
&lt;p&gt;Die Frage für Nutzer eingeschränkter Hardware ist nicht mehr: „Ist ein offenes Modell gut genug für die Aufgabe?&amp;quot; Die Frage veraltet monatlich. Hilfreicher ist: Kann diese Aufgabe bis morgens warten?&lt;/p&gt;
&lt;p&gt;Viele echte Aufgaben können das. Migrationen, Testabdeckung, Dependency-Updates, Dokumentation, Log-Analysen, der Longtail an Tickets, die niemand im Kopf behalten will. Diese Aufgaben brauchen Geduld, kein Tempo.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.stefanostermann.de/an-agent-while-i-sleep.html"/>
  </entry>
  <entry xml:lang="en">
    <title>An Agent That Works While I Sleep</title>
    <link rel="alternate" href="https://www.stefanostermann.de/an-agent-while-i-sleep.html"/>
    <id>https://www.stefanostermann.de/an-agent-while-i-sleep.html</id>
    <updated>2026-08-17T00:00:00Z</updated>
    <published>2026-08-17T00:00:00Z</published>
    <summary>The underrated advantage of running AI locally is not speed or quality. It is the eight hours when nobody has to watch.</summary>
    <content type="html">&lt;p&gt;Public AI has a billing meter. Every request adds up, and a long agent run in the evening eats the plan I still need for tomorrow&apos;s actual work. Local AI does not.&lt;/p&gt;
&lt;p&gt;Recently I used that difference on purpose. Before going to bed I handed my inference server the job I would not get done that evening: the missing CRUD endpoints in my current app, plus the tests that should have come with them. Clearing a field has to set it back to null, an organisation has to disappear along with its sites and appointments, and none of that was covered. Then I went to sleep.&lt;/p&gt;
&lt;h3&gt;After waking up&lt;/h3&gt;
&lt;p&gt;A branch with commits, the endpoints implemented on the server, and tests appended to each of the affected test files. A task list that started at one of five and ended empty. On llama.cpp with Qwen 3.8 27B.&lt;/p&gt;
&lt;p&gt;The interesting part is the price. The box pulls around 650 watts under load. Eight hours of that is roughly five kilowatt hours, which at my tariff is about €1.50 for a night of work. No token counter, no rate limit, no &amp;quot;you have reached your weekly usage&amp;quot; in the middle of a refactoring. I can start something that finishes in an hour and something that runs for three days.&lt;/p&gt;
&lt;h3&gt;The downsides&lt;/h3&gt;
&lt;p&gt;A finished branch does not mean finished. I read the diff in the morning, because an agent that had eight hours undisturbed also had eight hours to be confidently wrong. Slow review is, so to speak, the tax on unattended work.&lt;/p&gt;
&lt;p&gt;Two nights ago I woke up to a solution for a problem I described badly. The model had been extremely thorough in the wrong direction. That cost me a night.&lt;/p&gt;
&lt;p&gt;Locally on my notebook, the model is unbearably slow when I wait in real time. If I just go to bed, it is bearable!&lt;/p&gt;
&lt;h3&gt;Open questions&lt;/h3&gt;
&lt;p&gt;For people on limited hardware the question is no longer &amp;quot;is an open model good enough for the task?&amp;quot; That question gets older every month. The useful question is: can this task wait until morning?&lt;/p&gt;
&lt;p&gt;A lot of real work can. Migrations, test coverage, dependency updates, documentation, log analysis, the long tail of tickets nobody wants to hold in their head. Those tasks want patience, not speed.&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="de" href="https://www.stefanostermann.de/de/an-agent-while-i-sleep.html"/>
  </entry>
  <entry xml:lang="de">
    <title>Eine Woche Vibe Coding</title>
    <link rel="alternate" href="https://www.stefanostermann.de/de/vibe-coding.html"/>
    <id>https://www.stefanostermann.de/de/vibe-coding.html</id>
    <updated>2026-04-18T00:00:00Z</updated>
    <published>2026-04-18T00:00:00Z</published>
    <summary>Reflektionen über Produktivität, lokale LLMs und die Zukunft der Softwareentwicklung nach einer Woche &apos;Vibe Coding&apos;.</summary>
    <content type="html">&lt;p&gt;Nach einer Woche Vibe Coding nebenbei (während ich einen Fulltime-Job, Familie und viel zu wenig Schlaf jongliert habe), bin ich fast in ein richtiges Tief gerutscht. Dunkle Gedanken über die Zukunft meines Berufs und der Menschheit im Allgemeinen. Der Schlafmangel hatte wahrscheinlich auch etwas damit zu tun.&lt;/p&gt;
&lt;p&gt;Naja, ich habe ein paar Dinge gelernt...&lt;/p&gt;
&lt;h3&gt;Open-Source-Modelle&lt;/h3&gt;
&lt;p&gt;Sie sind immer noch nicht auf dem Niveau von Frontier-Modellen wie denen von Anthropic, aber die Fortschritte im letzten Jahr sind beeindruckend. Gemma 4 hat sich als besonders nützlich erwiesen und hat das meiste bewältigt, was ich ihm zugeworfen habe. Qwen 3.6 ist gerade erst erschienen, daher hatte ich noch nicht viel Zeit damit zu spielen. Manchmal ist eine zusätzliche Korrekturschleife im Vergleich zu Sonnet notwendig, aber ehrlich gesagt reicht es meistens aus. Meine größten Kopfschmerzen verursachten Bugs in den Modellen selbst oder in den Inference-Tools, die ich verwendet habe (llama.cpp-basiert, LM Studio).&lt;/p&gt;
&lt;h3&gt;Claude spielt in einer eigenen Liga&lt;/h3&gt;
&lt;p&gt;Ich habe Claude Code eine große, über 10 Jahre alte Codebasis hingeschmissen, und es hat nicht einmal gezuckt. In meiner Android-App (dem Versuchskaninchen für dieses Experiment) &lt;a href=&quot;https://www.hand-write.com&quot;&gt;HandWrite Pro&lt;/a&gt; habe ich eine Reihe neuer Funktionen, Stabilitäts-Fixes, UX-Verbesserungen, Modernisierungen und sogar einige architektonische Redesigns erstellt. Das alles als Nebenprojekt hätte normalerweise Monate, vielleicht Jahre gedauert.&lt;/p&gt;
&lt;p&gt;Anfangs lief alles reibungslos, es waren kaum Korrekturen nötig.&lt;/p&gt;
&lt;p&gt;Später zeigten sich erste Risse. Ein Refactoring führte zu einigen fiesen, schwer zu findenden Bugs. Und nachdem Sonnet und ich sie gemeinsam aufgespürt hatten, wurden dieselben Bugs wieder eingebaut. Und wieder.&lt;/p&gt;
&lt;p&gt;Das Schlimmste waren die Unit-Tests für eine Open-Source-PDF-Generierungsbibliothek. PDF-Generierung ist extrem komplex. Claude generierte zufrieden etwa 20 aufwendige PDF-Testdokumente mit allem Drum und Dran, aber jeder einzelne Test prüfte nur eine Sache: Ist die Datei größer als 0 Bytes? Das ist einfach faul. Fast schon Arbeitsverweigerung. Es brauchte viel gutes Zureden, bis Claude tatsächlich aussagekräftige Tests schrieb.&lt;/p&gt;
&lt;h3&gt;Fazit?&lt;/h3&gt;
&lt;p&gt;Open-Source-Modelle haben definitiv ihren Platz. Sie haben aufgeholt, sie sind solide und oft gut genug. Wann immer ich mit etwas persönlicheren Daten hantiere, verwende ich ein Open-Source-Modell.&lt;/p&gt;
&lt;p&gt;Frontier-Modelle wie Claude Opus oder Sonnet spielen in einer anderen Liga. Der Code, den sie produzieren, ist unglaublich gut. Jeder Entwickler, der sich weigert, mit LLMs zu arbeiten, wird zurückbleiben. Kein Mensch kann mit diesem Tempo mithalten.&lt;/p&gt;
&lt;p&gt;Aber sie sind nicht perfekt (das hat ja auch niemand erwartet, oder?). Was mich wahnsinnig machte, war, wie schnell ich an die Token-Limits stieß. Obwohl man auf dieses Problem natürlich Geld schmeißen kann. Das größere Problem ist wie verlockend es ist, das Gehirn einfach auszuschalten. Ich ertappte mich dabei, wie ich vom Sofa aus coden ließ, während ich eine Serie schaute, kaum aufpasste und einfach immer auf &amp;quot;OK&amp;quot; klickte, ohne mir die Shell-Befehle anzuschauen, die da ausgeführt werden sollten.&lt;/p&gt;
&lt;p&gt;Und ehrlich gesagt bleibe ich mit einem seltsamen Gefühl zurück, wohin das alles führt. Wie wird ein Informatikstudium in Zukunft überhaupt noch aussehen? Wie sollen wir unsere Kinder vorbereiten? Und wie vermeidet Europa es, ins Hintertreffen zu geraten, wenn alle guten LLMs aus den USA oder China kommen?&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.stefanostermann.de/vibe-coding.html"/>
  </entry>
  <entry xml:lang="en">
    <title>A Week of Vibe Coding</title>
    <link rel="alternate" href="https://www.stefanostermann.de/vibe-coding.html"/>
    <id>https://www.stefanostermann.de/vibe-coding.html</id>
    <updated>2026-04-18T00:00:00Z</updated>
    <published>2026-04-18T00:00:00Z</published>
    <summary>Reflections on productivity, local LLMs, and the future of software development after a week of &apos;vibe coding&apos;.</summary>
    <content type="html">&lt;p&gt;After a week of vibe coding on the side (while juggling a full-time job, family, and way too little sleep), I almost slipped into a proper funk. Dark thoughts about the future of my profession and humanity in general. Then again, the sleep deprivation probably had something to do with it.&lt;/p&gt;
&lt;p&gt;Anyway, I learned a few things...&lt;/p&gt;
&lt;h3&gt;Open source models&lt;/h3&gt;
&lt;p&gt;They&apos;re still not at the level of frontier models like Anthropic&apos;s, but the progress over the last year is impressive. Gemma 4 in particular turned out to be really useful and handled most of what I threw at it. Qwen 3.6 just dropped, so I haven&apos;t played with it much yet. Sometimes you need an extra correction loop compared to Sonnet, but honestly, it gets the job done. My biggest headaches were bugs in the models themselves or in the inference tools I was using (llama.cpp based, LM Studio).&lt;/p&gt;
&lt;h3&gt;Claude is on another level&lt;/h3&gt;
&lt;p&gt;I threw a large, 10+ year old codebase at Claude Code and it didn&apos;t even flinch. On my Android app (the guinea pig for this experiment) &lt;a href=&quot;https://www.hand-write.com&quot;&gt;HandWrite Pro&lt;/a&gt;, I created a bunch of new features, stability fixes, UX improvements, modernizations, and even some architectural redesigns. Doing all that as a side project would normally have taken me months, maybe years.&lt;/p&gt;
&lt;p&gt;At first everything ran smoothly with barely any corrections needed.&lt;/p&gt;
&lt;p&gt;Later on, cracks started to show. A refactoring introduced some nasty, hard-to-spot bugs. And after Sonnet and I tracked them down together, the same bugs got reintroduced. And again.&lt;/p&gt;
&lt;p&gt;The worst part was unit tests for an open source PDF generation lib. PDF generation is insanely complex. Claude happily generated about 20 elaborate PDF test documents with all the trimmings, but every single test only checked one thing: is the file bigger than 0 bytes? That&apos;s just lazy. Borderline refusing to work. It took a lot of nagging to get Claude to write actually meaningful tests.&lt;/p&gt;
&lt;h3&gt;So what do I take away from all this?&lt;/h3&gt;
&lt;p&gt;Open source models absolutely have their place. They&apos;ve caught up, they&apos;re solid, and they&apos;re often good enough. Whenever I am handling data that is at all personal, that is what I reach for.&lt;/p&gt;
&lt;p&gt;Frontier models like Claude Opus or Sonnet are in a different league. The code they produce is ridiculously good. Any developer who refuses to work with LLMs is going to get left behind. No human can keep up with that pace.&lt;/p&gt;
&lt;p&gt;But they&apos;re not perfect (nobody expected that anyway, right?). What drove me nuts was hitting token limits so quickly. Though I guess you can throw money at that problem. The bigger issue is how tempting it is to just switch your brain off. I caught myself coding from the couch while watching a show, barely paying attention and just clicking OK without looking at the shell commands that were about to run.&lt;/p&gt;
&lt;p&gt;And honestly, I&apos;m left with a weird feeling about where this is all going. What will a computer science degree even look like? How should we be teaching our kids? And how does Europe avoid getting left in the dust when all the good LLMs come from the US or China?&lt;/p&gt;
</content>
    <xhtml:link rel="alternate" hreflang="de" href="https://www.stefanostermann.de/de/vibe-coding.html"/>
  </entry>
</feed>
