Working notes on building with LLMs — patterns, harnesses, evals, the parts nobody puts in the demo. Written primarily for future-me opening a fresh chat. You’re welcome to read along.
I went looking for the next big model drop for the week of Aug 4-10 and found a quiet flagship front — one safety system shipped, no benchmark leaps.
No new flagship — OpenAI tiered GPT-5.6 by price and risk, Neon and Castform shipped a 100x cheaper retriever, and Anthropic put Claude Code on autopilot.
Flagships were quiet Aug 4-10. Open source wasn't — a 30B local agentic model under Apache 2.0, a 2.6B on-device runner, and a wave of agent memory tooling.
EU transparency rules went live, India set a 3-hour takedown clock, California launched an AI cyber program, and Google chased a $1.5B coding agent deal — a busy week for AI regulation and industry moves.
No flagship model shipped — instead we got routers that cut costs 68%, MCP growing up with a Ruby SDK and an Agents working group, and coding tools turning into model supermarkets.
No flagship model dropped July 28 to August 3 — the labs spent the week slicing price and tiering access instead of chasing a bigger benchmark number.
One lab shipped three products in one week — none of them chatbots. And enterprise adoption was all partnerships, no deployment numbers.
July 28 to Aug 3 was quiet for open weights and papers — the only verifiable open source ship was MCP going stateless on July 28, plus DeepSeek's V4 Flash hitting beta July 31.
The EU's GPAI rules went enforceable August 2 while Washington offered no new AI law — a quiet U.S. week that says as much about direction as any executive order.
GitHub didn't ship a new model July 28–Aug 3 — it shipped the plumbing: review that runs your tools, PRs that stack, and agents you trigger with a comment.
A quiet flagship week where Anthropic owned the window with Claude Opus 5 — not a raw intelligence leap, but a cost-per-task one.
One 2.8T open model shipped with day-0 infra from vLLM and Telnyx. Anthropic bought enterprise distribution via Cognizant instead.
For July 21-27 open source shipped the only frontier-scale drop of the week — a 2.8T Kimi K3 — plus the RL infra that trained it, a real-time voice runtime, and four papers that fix boring but real problems.
For July 21-27, AI regulation mostly held its breath — no new EU enforcement, no new U.S. law, and industry kept its checkbook closed before the August rush.
For July 14-20, no lab shipped a new frontier model — DeepMind's Gemini 3.5 Flash Cyber was the only benchmarked release, and it was built to find vulns, not win MMLU.
No flagship model shipped. Enterprises bought quality gates and cost routers — plumbing laid the week before open weights jumped an order of magnitude.
July 14-20 had no flagship fireworks. Open source did the real shipping — a 975B MoE from Thinking Machines, a run of tiny specialist embeds, and papers that poke holes in how we eval agents.
For July 14-20, Brussels ordered Google to open Android and Search to rival AI assistants, set the clock for AI label rules on August 2, and left US states and the labs themselves to fill a federal vacuum.
July 7-13 was loud: OpenAI shipped GPT-5.6 and xAI shipped Grok 4.5 within a day of each other, while the only clean benchmark in the window came from Mistral's 8B Robostral Navigate.
The week flagships actually shipped — GPT-5.6 and Grok 4 on the same day, GPT-Live the day before, and three OpenAI enterprise stories that matter less than the Copilot distribution deal.
July 7-13, open source skipped the chatbot leaderboard and shipped embodied AI, medical data plumbing, and a sovereign European MoE.
For July 7-13, the FTC said hidden output steering could be deception, Brussels gave a final green light to its AI Act Omnibus, and OpenAI's next frontier rollout reportedly routed through a White House review.
June 30 to July 6 gave us a cheaper Sonnet that chases Opus, a 6B-active prover that saturates miniF2F, and a new bio benchmark with no scores yet.
June 30 to July 6 didn't ship a frontier model. It shipped the deployment orgs — AWS and Microsoft committed $3.5B to embed engineers, plus a science workbench and cheaper plumbing from Google and X.
June 30-July 6 had no new open flagship, but it shipped a compiler that turns prompts into local adapters, a 6T-token VLM data playground, and a world model that freezes itself.
For June 30 to July 6, the FTC framed hidden AI steering as deception, Brussels formally adopted its AI Act omnibus, and AWS bet $1B that deployment — not training — is the moat.
June 23-29 had no open flagship launch — instead GPT-5.6 was throttled at the White House's request and Mythos went gated to 100+ orgs.
Frontier models went gated by government, Ford brought back 350 veteran engineers after AI fell short, and Google made Gemini personalization free — the same week South Korea bet $1T on physical AI.
No open flagship shipped June 23-29. Open source did — four verified drops from 0.6B to 75B, all about making models cheaper and faster to run.
For June 23-29, Washington pulled the brake pedal for robotaxis while tightening the gates on chips and frontier models — and Brussels mostly waited for August.
For June 16-22 I couldn't verify a single flagship model launch or benchmark jump — the labs shipped scaffolding, not new weights.
June 16-22 didn't ship a flagship model. It shipped the surfaces around it — Gemini on your shelf and in your pocket, ChatGPT on Samsung devices, and a security platform designed to make enterprise adoption boring.
June 16-22 was long-horizon week in open source — a 1M-token GLM-5.2 under MIT, Baidu's constant-memory OCR, and agents that learn to scaffold themselves.
June 16-22 was quiet on formal AI law — no FTC action, no EU text, no Hill vote I could verify. The policy moved anyway through enterprise contracts and security packaging.
For June 9-15, one real frontier model shipped — Claude Fable 5 and Mythos 5 — then got pulled three days later. Google filled the gap with speech and speed, not a Gemini flagship.
No flagship model June 9-15. OpenAI shipped where agents run (Ona), where enterprises buy (Oracle + $150M partner network), and BBVA proved it scales to 100,000 seats.
June 9-15, the only flagship that shipped got pulled. Open source shipped the fix for 1M-token context — two sparse attention systems, a 30B MoE that actually does video, and a way to compress agent skills.
For June 9-15, Washington told Anthropic to pull its most capable models worldwide, states moved to fill the federal vacuum, and the security community said the ban hurts defenders more than adversaries.
For June 2-8, no flagship model shipped and no benchmark moved — the labs were heads-down before Fable 5 landed the next day.
No flagship product launched June 2-8. One preview shipped, WWDC opened, and everything else was heads-down before the enterprise push that landed June 9.
No open model dropped June 2-8. One NVIDIA paper did — SparDA — and it's the clearest hint at how 1M context gets cheap.
For June 2-8, no AI law passed and no regulator moved — but GitHub shipped governance plumbing and WWDC opened, right before export bans rewrote the rules the next week.
For May 26 to June 1 I looked for a new flagship or a benchmark move and found neither — the labs were quiet the week before the June lull.
No flagship model dropped May 26–June 1 — instead the labs shipped enterprise packaging: controls, connectors, and distribution deals that decide who actually gets to use the models.
May 26-June 1, open source didn't chase scale — it chased efficiency. Two MoEs that activate less, a vision model with no encoder, and two systems that turn video into action.
For May 26–June 1, no AI law passed in Washington or Brussels — the policy moved through admin controls, marketplace deals, and procurement instead.
May 19–25 was a quiet week at the frontier — no verified flagship model release and no new benchmark card to score.
May 19–25 didn't ship a flagship model. It shipped the unsexy stuff that decides enterprise adoption anyway — admin controls, cost scrutiny, and trust boundaries.
May 19-25 had no flagship model drop — instead, open source shipped efficiency. A $1,500 pretrain, a 3.8B image model on 19% compute, and RLVR that finally cares about tokens.
No AI Act dropped May 19-25, but two executive orders rewrote the playbook for AI in finance and NIST put $20M per pilot behind AI on the factory floor.
A growing share of traffic to technical blogs comes from agents fetching pages to answer someone else's question. Writing for that reader doesn't mean robotic prose — it means structuring posts so the load-bearing claims survive extraction.
After a year of building agents at Axonify, the thing that actually changed how my team ships isn't a smarter model — it's the Model Context Protocol. Here's why.
Prompt engineering was a useful crutch when models were smaller and more brittle. With today's frontier models and large context windows, the real lever is what you put in the context, not how you phrase it.
The 'AI replaces developers' framing is bad for both your career and your codebase. After two years of agent-assisted shipping at Axonify, here's what actually works.