Openrouter
3 posts
-
An early-access test of TypeSafe's Jev: calibrated judgments for half a cent
An early-access test of TypeSafe's Jev, a model that answers typed questions with probabilities and writes no text. I ran it on 24 Norwegian responses to the 2022 hearing on resource rent tax for salmon farming. Agreement with a frontier model's labels, calibration, cost and latency, next to DeepSeek V4.1 Flash with and without reasoning, and why an evolutionary economist goes looking for the odd variant.
-
Measuring what a coding agent actually costs
A self-hosted OpenTelemetry stack for LLM usage, built as an OpenObserve experiment for customers who keep telemetry on prem: two months of getting the numbers wrong, 118,469 spans of which 118 were useful, and what Rosenberg's learning by using says about why the price list told me nothing.
-
Forking Codex to talk to any endpoint
How we restored the Chat Completions wire API in a Codex fork, so one coding agent drives local Ollama models, OpenRouter, Azure OpenAI with Entra auth, or any EU-hosted provider. The three patches that made it usable, the branch model that keeps the rebases to an hour, and why Baldwin & Clark's design rules say the interface was the thing to fight for.