We Asked for Metrics. We Got Attribution Theater.

There is a pattern in senior resumes and LinkedIn profiles that makes me more uncomfortable every year. Someone “grew revenue by 38%.” Someone else “reduced costs by $4.2M.” Another “increased engineering productivity by 60%.” Sometimes I know the work behind these lines well enough to know that the link between the person and the number was much weaker than the sentence suggests. It is tempting to call this a problem of dishonest people. My view is that it is worse than that. I believe we built a market that rewards the performance and punishes the truth. In my experience, the candidate who writes “grew revenue by 38%” is more likely to get the interview than the one who honestly describes a shared effort. When the reward for exaggeration is high and the chance of being checked is low, I do not see exaggeration as a personal failure. I see it as the rational response to the system we built. ...

September 19, 2026 · 7 min · Rami Pinku

Judgment at the Frontier

Dario Amodei, CEO of Anthropic, has published an unusually direct argument about the future of AI development: the companies building the most capable models need to slow down. He is not calling for them to stop, or to give up the benefits AI may bring. He is asking them to leave enough time to understand, evaluate, and control the systems they are making more powerful. What makes this worth reading is who is saying it. Amodei runs one of the companies competing to build those models. He faces the same pressure to move faster as everyone else, and he is arguing that the pace itself has become part of the problem. ...

September 13, 2026 · 6 min · Rami Pinku

Anthropic Just Described the Operating Model I've Been Writing About for a Year

Anthropic published a new piece last week, The AI-Native SDLC Playbook. It is long, detailed, and worth the time. It also lines up closely with what I have been writing here for the better part of a year. The argument at its center takes one sentence: code is no longer the bottleneck. Everything in the document follows from it. Once agents write most of the implementation, the constraint moves to the stages on either side, the ones still running at human speed. Approval queues build. Controls sized for human output stop matching what arrives at them. Governance that meets monthly cannot govern work that ships hourly. ...

August 29, 2026 · 2 min · Rami Pinku

The Skill Records Its Reasoning. It Does Not Record Yours.

On a recent episode of Aakash Gupta’s podcast, Oji Udezue ran a product idea through a viability gate he had built as a Claude Code skill. The idea was a tool called Standup Zero that reads Slack comments and assembles a daily standup digest. The gate scored it on six dimensions and returned three moderates. Problem clarity and urgency, moderate, because workflow convenience is not a deep problem. Target user definition, not strong. Differentiation, not strong. Competitive landscape, strong, which on this scorecard counts against you, because it means everyone can build what you are building. ...

August 8, 2026 · 6 min · Rami Pinku

Fail Fast Was Never a Strategy.

A June 2026 NBER working paper matched more than 100,000 GitHub developers against their AI usage telemetry. Developers using autonomous agents committed 180 percent more code. Their releases rose 30 percent. The gain is enormous where code gets written, and it decays at every step after that. Marty Cagan named the phenomenon last week: the AI productivity paradox. He also points out that while nearly everyone now agrees it exists, almost nobody agrees on why. ...

July 25, 2026 · 5 min · Rami Pinku

The Right Tool, Not the Newest One

I recently needed to understand a long, cross-departmental process. It was documented across a huge slide deck, a few forms, and half a dozen fields scattered across internal systems. No single source of truth. In my experience, the right tool for this is a checklist. Not a dashboard, not an assistant. Just a simple list of boxes to check in the right order. So I used Claude to read through the materials and produce a tight summary of the process and its steps. I reviewed it, then asked Claude to turn that summary into a checklist draft. I refined it. I’m using it now, and I’ll keep refining it as I go. ...

July 18, 2026 · 4 min · Rami Pinku

The Judgment Log at the Engineering Station: What to Write, What to Skip

The last post left the engineering station with a one-line test: if the reasoning behind a decision would be useful to the engineer who touches this code in six months, write it down. That test tells you what deserves an entry, but it does not tell you how to recognize the moment, in the middle of a session, when you have crossed from accepting a suggestion into making a decision. Those are different skills, and the second one is a skill nobody ever had to develop before. ...

July 11, 2026 · 6 min · Rami Pinku

I Spent $350 to Learn Elena Verna Is Right

A few days ago I kicked off several research tasks in parallel and didn’t check which model was running underneath. All of them defaulted to Claude Opus 4.8. When I looked at the bill, I’d spent $350. So the question became obvious. Was the output better? No. I compared the results against similar research I’d run before on cheaper models. No meaningful difference in quality. I still had to verify every source and go deeper on half the areas myself. The expensive model just sounded more confident while doing it. ...

July 4, 2026 · 4 min · Rami Pinku

The Judgment Log in Practice: One Chain, Four Stations

I ended the last post with a question: the next challenge is not building the Judgment Log. It is whether anyone writes in it once the deadline is two hours away. That question only has a useful answer if the artifact is light enough to actually use. So instead of arguing for it further, I want to show it. Take a fictional but familiar scenario. A checkout flow. A promotional window. A promo code validation feature built using AI-assisted development. It shipped. Three weeks later, it broke when the campaign introduced expired codes. In the post-mortem, nobody could answer the three questions that mattered: what did the PM cut and why, what did the designer choose between, and what did the engineer override. ...

June 27, 2026 · 7 min · Rami Pinku

The Judgment Log: The Artifact JDD Teams Need

In April, Meta employees burned through 73.7 trillion tokens in roughly thirty days. The company found out not because spending crossed some alarming threshold, but because an internal leaderboard, nicknamed Claudeonomics, had turned token consumption into a competition. Employees and teams were ranked by how much they used. The system did exactly what it was built to do: usage went up. What it could never show anyone was whether any of that usage produced something worth the cost. Meta is now dismantling the leaderboard in favor of a centralized monitoring platform called AI Gateway, built to track spending in real time and flag unusual spikes. ...

June 20, 2026 · 7 min · Rami Pinku