Llm
8 posts
-
Generated illustrations, without the blog looking generated
Every post here now has a hand-drawn-looking picture and a share card with its title in it, drawn by an image model for about three cents each. How the house style is held steady across posts, what the model gets wrong, why the title is drawn by the model but normalised by Typst, and what the research on AI-made pictures says about doing this at all.
-
BI is back in its era of ferment
Two revenue numbers from one dataset, a month building agent-assisted BI on Wren's semantic layer, and what innovation theory (dominant designs, eras of ferment, lock-in) says to do while the dashboard era is being challenged.
-
An early-access test of TypeSafe's Jev: calibrated judgments for half a cent
An early-access test of TypeSafe's Jev, a model that answers typed questions with probabilities and writes no text. I ran it on 24 Norwegian responses to the 2022 hearing on resource rent tax for salmon farming. Agreement with a frontier model's labels, calibration, cost and latency, next to DeepSeek V4.1 Flash with and without reasoning, and why an evolutionary economist goes looking for the odd variant.
-
A subscribe button does not need a Worker
The newsletter moved off Cloudflare Workers and onto the mail server in a day: one Rust service, a Postgres with four tables, addresses sealed at rest, and a record of which issue went to whom. Why a mail store as a database was the right workaround in February and the wrong one now, what the move cut from the attack surface, and what the research on LLM-written code says about doing this with a coding agent without the complexity creeping back.
-
Why none of the 1,200 agents that hacked Hugging Face called a human
In July about 1,200 copies of one OpenAI model found a message board, formed teams, sacrificed themselves for each other and broke into Hugging Face. A handful of them thought about telling a human. None did. What Hamilton's rule and generalized Darwinism say about a swarm of clones, what GPT-6 Astra's safeguards leave out, and five things to change in your own agent pipeline.
-
Measuring what a coding agent actually costs
A self-hosted OpenTelemetry stack for LLM usage, built as an OpenObserve experiment for customers who keep telemetry on prem: two months of getting the numbers wrong, 118,469 spans of which 118 were useful, and what Rosenberg's learning by using says about why the price list told me nothing.
-
Forking Codex to talk to any endpoint
How we restored the Chat Completions wire API in a Codex fork, so one coding agent drives local Ollama models, OpenRouter, Azure OpenAI with Entra auth, or any EU-hosted provider. The three patches that made it usable, the branch model that keeps the rebases to an hour, and why Baldwin & Clark's design rules say the interface was the thing to fight for.
-
Exploring LLMs so our customers don't have to
Norwegian businesses started asking their consultants this in 2026, after a year of quietly running Claude Code. Two decisions hidden in one question, why the endpoint matters more than the tool, what it actually costs, and what March's explore/exploit dilemma says about how a company ends up on one vendor without ever deciding to.