<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
    <channel>
      <title>John Wilger</title>
      <link>https://johnwilger.com</link>
      <description>Senior Engineering Leader &amp; Principal Engineering Consultant</description>
      <generator>Zola</generator>
      <language>en</language>
      <atom:link href="https://johnwilger.com/rss.xml" rel="self" type="application/rss+xml"/>
      <lastBuildDate>Sun, 26 Jul 2026 00:00:00 +0000</lastBuildDate>
      <item>
          <title>Somebody Has to Design the Factory</title>
          <pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/somebody-has-to-design-the-factory/</link>
          <guid>https://johnwilger.com/blog/somebody-has-to-design-the-factory/</guid>
          <description xml:base="https://johnwilger.com/blog/somebody-has-to-design-the-factory/">&lt;p&gt;Four rules got added to my agent harness &lt;a href=&quot;&#x2F;blog&#x2F;the-software-i-wouldnt-have-bothered-to-build&#x2F;&quot;&gt;the morning I built
Lanyard&lt;&#x2F;a&gt;. Ticket work has to
happen in an isolated worktree. Commits and pushes from the primary checkout are
blocked. A deployment isn&#x27;t a deployment until a live URL answers. And a
task-closing command&#x27;s exit code doesn&#x27;t count as evidence that the task closed.&lt;&#x2F;p&gt;
&lt;p&gt;Each of them came from something going wrong that day. None of them are features
of Lanyard, which is an SSH proxy and has nothing to do with any of it. Lanyard
was close to a byproduct. What I actually spent the morning on was a slightly
better factory.&lt;&#x2F;p&gt;
&lt;p&gt;That word has started showing up in conversations about where agentic development
is heading, usually as &quot;dark factory,&quot; and about half the time I use it I can
watch the person across from me hear something I didn&#x27;t say.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-people-picture&quot;&gt;What people picture&lt;&#x2F;h2&gt;
&lt;p&gt;The image is a text box. You type what you want, some quantity of machine
intelligence happens, and a finished product comes out the other side. No
engineers, no designers, no product manager, no argument about what the thing
should do. The prompt is the spec, the output is the product, and the entire
middle got vaporized.&lt;&#x2F;p&gt;
&lt;p&gt;Products shaped like that exist and some are good at what they do, in exactly the
way Canva is good at making you a &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;www.canva.com&#x2F;newsroom&#x2F;news&#x2F;brand-campaign-2026&#x2F;&quot;&gt;coffee mug with a squirrel on
it&lt;&#x2F;a&gt;. They aren&#x27;t what
the dark factory argument is about, and treating them as its endpoint makes the
actual claim impossible to discuss.&lt;&#x2F;p&gt;
&lt;p&gt;What&#x27;s happening there is a category error rather than a disagreement. It
collapses three separate activities into one: deciding what to build, engineering
how it works, and producing it. A dark factory isn&#x27;t a wish granted. &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;www.industrialautomationindia.in&#x2F;articles&#x2F;fanuc-lights-out-manufacturing-benchmark&quot;&gt;FANUC has
run one in Oshino since
2001&lt;&#x2F;a&gt;,
where robots assemble roughly fifty robots a day and the line runs unsupervised
for up to thirty days at a stretch. The lights are off because nobody is on the
floor to need them. According to one of their VPs, they turn off the heat and air
conditioning too.&lt;&#x2F;p&gt;
&lt;p&gt;That plant is one of the most heavily engineered objects on the planet. Every one
of those thirty unattended days rests on people who decided what a robot should
be, people who worked out how to make one, and people who worked out how to make
a machine make one. Nobody typed &quot;robot&quot; into anything.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;benz-s-workshop-had-no-factory&quot;&gt;Benz&#x27;s workshop had no factory&lt;&#x2F;h2&gt;
&lt;p&gt;When Carl Benz and Gottlieb Daimler were building the first automobiles in the
mid-1880s, they worked in engineering workshops. Parts were forged, cast, and
machined, often by specialist suppliers who had been making something else
entirely the year before. Then skilled workers assembled the car by hand, filing,
shimming, and reaming each piece until it mated with the one next to it.&lt;&#x2F;p&gt;
&lt;p&gt;That fitting wasn&#x27;t unskilled labor. A fitter was reading the actual assembly in
front of him, feeling where it bound, and correcting for the gap between what the
drawing claimed and what the shop had produced.&lt;&#x2F;p&gt;
&lt;p&gt;An enormous amount of product design and mechanical engineering went into those
cars, and much of the learning was only available with your hands on the thing.
You found out how a component behaved by making one, setting it next to another
one, and watching what happened.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;ford-automated-the-fitting-not-the-designing&quot;&gt;Ford automated the fitting, not the designing&lt;&#x2F;h2&gt;
&lt;p&gt;The moving assembly line arrived at Highland Park in 1913, and the thing that
made it possible wasn&#x27;t the conveyor. It was parts held to tolerance tightly
enough that fitting became unnecessary. A line only works if the part coming down
the chute drops into place without anyone deciding anything about it. Ford didn&#x27;t
invent that idea; armory practice in the American arms industry had been chasing
interchangeable parts for most of the preceding century. Ford industrialized it,
at a scale and a price nobody had managed before.&lt;&#x2F;p&gt;
&lt;p&gt;Getting there took a body of engineering that had barely existed as a job. Somebody
had to design the gauges that said whether a part was in spec. Somebody had to
design the jigs and fixtures that held the work in exactly one orientation so it
couldn&#x27;t be machined backwards. Somebody had to cut the dies, balance the stations
so no position starved or flooded the next, and run the tolerance stack-up analysis
that tells you whether a chain of individually acceptable parts still adds up to a
door that closes.&lt;&#x2F;p&gt;
&lt;p&gt;Product design didn&#x27;t shrink when the line showed up. A second design discipline
grew up beside it, hard and expert and well paid, and we now call it manufacturing
engineering.&lt;&#x2F;p&gt;
&lt;p&gt;Notice what the tolerance regime actually eliminated. The fitter with a file was a
human compensating, one part at a time, for a process that couldn&#x27;t hold its own
spec. Holding the spec is what made the compensation unnecessary. Filing the part
down is what you do when the factory doesn&#x27;t work.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-fitters-didn-t-become-tool-designers&quot;&gt;The fitters didn&#x27;t become tool designers&lt;&#x2F;h2&gt;
&lt;p&gt;The aggregate story is that automotive engineering employment grew enormously over
the following century. The story for an individual fitter in 1913 was worse than
that.&lt;&#x2F;p&gt;
&lt;p&gt;Work on the line was miserable in a way the craft work hadn&#x27;t been. Ford hired
50,448 people that year to sustain an average workforce of 13,623, a turnover rate
around 370 percent, which meant that growing the workforce by a hundred took
hiring nearly a thousand. In January 1914 Ford answered by roughly doubling pay:
&lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;www.thehenryford.org&#x2F;collections&#x2F;explore&#x2F;articles&#x2F;fords-five-dollar-day&quot;&gt;five dollars for an eight-hour
day&lt;&#x2F;a&gt;
against the previous $2.34 for nine. Turnover fell to 54 percent. Nobody doubles
wages to keep people in a job they like.&lt;&#x2F;p&gt;
&lt;p&gt;So the claim I&#x27;m making is narrow. The work didn&#x27;t disappear, and there turned out
to be more of it, at higher skill, than before. That isn&#x27;t the same statement as
&quot;your job is safe,&quot; and I&#x27;m not making the second one. Aggregate growth is cold
comfort when the growth lands on different people, in a different city, a decade
later. Anyone promising you that the transition to agentic development will be
smooth for everyone currently employed is guessing, and guessing in the direction
that happens to be comfortable for them.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;nobody-sells-you-a-factory&quot;&gt;Nobody sells you a factory&lt;&#x2F;h2&gt;
&lt;p&gt;Cars didn&#x27;t get simpler. Neither did the engineering. A modern vehicle program
employs more engineers across more disciplines than Benz could have imagined, and
it delivers more variety at lower cost to more people.&lt;&#x2F;p&gt;
&lt;p&gt;What matters most for software is what happens when a company decides to build a
new model. They don&#x27;t roll the last one&#x27;s factory over onto it.
They design a production system for that product: new dies, new fixtures, new
gauges, a rebalanced line, a revalidated supply chain. Platforms and plants get
shared, and a flexible line can build more than one model, but the tooling specific
to a product is designed fresh, by experts, at large expense, every single program.&lt;&#x2F;p&gt;
&lt;p&gt;That&#x27;s the fact that should make you suspicious of anyone offering to sell you a
software factory. In manufacturing, machinery is a thing you can buy. The
production system is a thing your engineers design for what you&#x27;re building. When
somebody&#x27;s pitch is that you can buy the second one, they&#x27;re selling you machinery
and calling it a factory.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-line-produces-changes-not-products&quot;&gt;The line produces changes, not products&lt;&#x2F;h2&gt;
&lt;p&gt;The analogy has an obvious hole and it&#x27;s worth walking straight into it. Physical
factories exist because replication is expensive. You design a car once and then
face the separate, enormous problem of producing a hundred thousand identical ones.
Software solved that problem decades ago. Copying a build costs nothing. &lt;code&gt;cp&lt;&#x2F;code&gt; is
the entire assembly line, and it was never the interesting part.&lt;&#x2F;p&gt;
&lt;p&gt;So what would a software factory be for?&lt;&#x2F;p&gt;
&lt;p&gt;The thing that repeats in software isn&#x27;t the product, it&#x27;s the process. The unit
coming off the line is a change: a feature, a fix, a migration, a dependency bump.
Each one runs the same sequence. Work out what the behavior should be, decide what
will be expensive to reverse later, write a test that fails for the right reason,
make it pass, review it, integrate it, produce evidence that it actually deployed.
Different content every run, identical sequence.&lt;&#x2F;p&gt;
&lt;div class=&quot;diagram-scroll&quot; role=&quot;region&quot; aria-label=&quot;Scrollable diagram: different software changes move through one validated sequence&quot; tabindex=&quot;0&quot;&gt;
  &lt;img src=&quot;&#x2F;images&#x2F;blog&#x2F;somebody-has-to-design-the-factory-change-line.svg&quot; alt=&quot;A feature, fix, migration, and dependency bump entering the same sequence of behavior definition, reversal-risk analysis, a failing test, implementation, review, integration, and live deployment evidence.&quot;&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;A factory that outputs changes is still a factory, and it describes what&#x27;s being
automated better than a conveyor of identical bodies does. The closer physical
analogue is a job shop running standard work: every job different, every job pushed
through the same validated process with the same fixtures and the same gauges.&lt;&#x2F;p&gt;
&lt;p&gt;This is where the analogy strains hardest. A stamping press doesn&#x27;t read an
acceptance criterion and decide whether it&#x27;s been met. My harness has to. The
automated steps need judgment in them, and bounding that judgment is the whole
design problem: enough to absorb a different change every time, not so much that
the process becomes a negotiation.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;tolerances-for-software&quot;&gt;Tolerances, for software&lt;&#x2F;h2&gt;
&lt;p&gt;Tolerance stops being a metaphor as soon as you ask what the harness refuses to do.&lt;&#x2F;p&gt;
&lt;p&gt;Production code can&#x27;t exist before a failing test has been observed, and the
failure has to come from the behavior being missing rather than from a typo or a
broken fixture. Code that shows up without that gets reverted rather than debated.
That&#x27;s a station that can&#x27;t be run out of sequence.&lt;&#x2F;p&gt;
&lt;p&gt;Test evidence is bound to a hash of the exact diff it was produced against, and the
clean-review streak resets to zero the moment that hash changes. Three clean review
passes followed by one more edit is zero clean review passes. An inspection record
belongs to a specific part, not to the general vicinity of the part.&lt;&#x2F;p&gt;
&lt;p&gt;A critical security or human-safety problem introduced by the change itself can
only be dispositioned as fixed. &quot;Defended&quot; and &quot;accepted risk&quot; exist for other
categories and are unavailable for that one. That&#x27;s a gauge rejecting a part rather
than a reviewer holding an opinion about it.&lt;&#x2F;p&gt;
&lt;p&gt;CI recovery runs off a table compiled into the tool instead of described in a
document. A failure classified as caused by the change may only be repaired. One
classified unrelated or transient may only be rerun at the unchanged commit. There
is no third option, the classification has to be recorded before an action is
chosen, and the hold doesn&#x27;t release until a matching replacement run reaches
terminal success. Queued isn&#x27;t success. Neither is a zero exit code.&lt;&#x2F;p&gt;
&lt;p&gt;Reviews run in fresh contexts that have to attest they were fresh, and the
reviewing role has no write tools at all, because an inspector who can adjust the
part isn&#x27;t an inspector.&lt;&#x2F;p&gt;
&lt;p&gt;The rule underneath all of them is the one I&#x27;d point at first. No agent working in
my repositories may install a substitute tool, weaken a required gate, or fall back
on stale memory in order to keep moving. When nothing available can satisfy a gate,
the correct behavior is to stop at the gate and come get me.&lt;&#x2F;p&gt;
&lt;p&gt;That rule is the fitter&#x27;s file, confiscated on purpose. Weakening a gate so a
change can land is hand-fitting: a human compensating, one part at a time, for a
process that can&#x27;t hold its own spec. It works, it&#x27;s invisible in the finished
product, and it&#x27;s the reason nothing downstream of it can be trusted. The entire
point of a tolerance regime is that nobody has to reach for the file.&lt;&#x2F;p&gt;
&lt;img src=&quot;&#x2F;images&#x2F;blog&#x2F;somebody-has-to-design-the-factory-fitter-and-gauge.png&quot; alt=&quot;On the left, a worker files one mismatched component until it fits; on the right, a fixed gauge rejects an out-of-tolerance component before it enters an automated line.&quot; loading=&quot;lazy&quot;&gt;
&lt;p&gt;Stations get different capability on purpose. Bounded mechanical work runs on a
small model with read-only tools and has to hand back evidence the caller can
verify independently, because its own account of what it did doesn&#x27;t count.
Architecture and security analysis go to a stronger model that also can&#x27;t write.
That&#x27;s a fixture-and-gauge argument rather than a model-quality argument.&lt;&#x2F;p&gt;
&lt;p&gt;None of this is the model getting better. I&#x27;ve written before about &lt;a href=&quot;&#x2F;blog&#x2F;make-fabrication-unrepresentable&#x2F;&quot;&gt;making bad
states structurally impossible rather than
discouraged&lt;&#x2F;a&gt;, and about how &lt;a href=&quot;&#x2F;blog&#x2F;coding-harness-plugins-are-agentic-systems&#x2F;&quot;&gt;a guardrail
nobody has tested for behavior change may not be a
guardrail&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-part-that-needs-your-best-people&quot;&gt;The part that needs your best people&lt;&#x2F;h2&gt;
&lt;p&gt;The product manager still decides what the product is. The senior engineers still
decide which parts go together and which choices will be expensive to reverse.
None of that moved. What&#x27;s new is that those same people are needed to design the
factory, and that&#x27;s the work I&#x27;d put the most experienced person in the room on.&lt;&#x2F;p&gt;
&lt;p&gt;Every gate I described started as something a staff-level engineer already knows.
Don&#x27;t trust an exit code. Don&#x27;t let a green review outlive the diff it reviewed.
Don&#x27;t let the author sign off on their own work. That knowledge is normally tacit,
exercised a few dozen times a week by people who would struggle to write it down if
you asked them to.&lt;&#x2F;p&gt;
&lt;p&gt;Encoding it is harder than exercising it. Exercising judgment lets it stay vague.
Encoding it forces you to state exactly what the standard is, exactly what evidence
satisfies it, and exactly what happens when the evidence is missing, in a form
something else can check with you out of the room. Most teams have never had to do
that, which is a large part of why they don&#x27;t have a factory. The machinery isn&#x27;t
what&#x27;s missing. What&#x27;s missing is a written-down definition of good, precise enough
to gauge.&lt;&#x2F;p&gt;
&lt;p&gt;Two claims get mashed together here and they&#x27;re worth separating. The line didn&#x27;t
make any individual worker faster at fitting a part. It removed the fitting. What
improved was throughput and rework, which are properties of a system rather than of
anybody&#x27;s hands. I&#x27;ve argued that &lt;a href=&quot;&#x2F;blog&#x2F;ai-might-not-make-you-faster-it-can-make-you-free&#x2F;&quot;&gt;AI probably isn&#x27;t making you
faster&lt;&#x2F;a&gt;, and I still
think that holds for one person doing one task. A factory is a different unit of
analysis. You can have a well-designed one and remain personally slower at
everything, and that would be fine.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;where-the-lights-stay-on&quot;&gt;Where the lights stay on&lt;&#x2F;h2&gt;
&lt;p&gt;The lights go off over the line. They stay on in the room where somebody decides
what to build, and in the room next door where somebody decides how it gets built
the same way ten thousand times.&lt;&#x2F;p&gt;
&lt;p&gt;That second room barely existed for software five years ago. It&#x27;s where I&#x27;ve spent
most of the last year.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>The Software I Wouldn&#x27;t Have Bothered to Build</title>
          <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/the-software-i-wouldnt-have-bothered-to-build/</link>
          <guid>https://johnwilger.com/blog/the-software-i-wouldnt-have-bothered-to-build/</guid>
          <description xml:base="https://johnwilger.com/blog/the-software-i-wouldnt-have-bothered-to-build/">&lt;p&gt;This morning I had an annoyance. By lunchtime it had a name, an architecture,
a public repository, a Rust implementation, ADRs, strict tests, release
automation, and &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;jwilger.github.io&#x2F;lanyard-ssh-agent&#x2F;&quot;&gt;a website that looks better than most developer-tool sites
have any right to&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The tool is &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;lanyard-ssh-agent&quot;&gt;Lanyard&lt;&#x2F;a&gt;, an SSH-agent
switching proxy. I built it because I use 1Password&#x27;s SSH agent when I&#x27;m sitting
at my desktop, but I also SSH into that desktop from my laptop and forward the
laptop&#x27;s agent. Both routes work fine until they meet a long-lived Zellij
session.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;code&gt;SSH_AUTH_SOCK&lt;&#x2F;code&gt; inside the session points at whichever socket existed when the
session started. Attach locally and the session knows about 1Password. Attach
remotely and it may know about the forwarded agent. Move between the two and
the environment doesn&#x27;t follow you. Authentication and Git commit signing then
depend on which machine happened to create the session hours or days earlier.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;d put up with this for a while. The usual next step would have been a shell
function that rewrote an environment variable, a symlink shuffled between
sockets, or some startup hook whose behavior I&#x27;d forget six months later. It
would mostly work. It would also fail in exactly the cases that made the problem
interesting.&lt;&#x2F;p&gt;
&lt;p&gt;Instead, I spent some time defining the problem with a coding agent. We worked
through the architecture, wrote down the decisions that would be expensive to
change later, and split the work into increments I could review independently.
Then I left the computer and did other things with my life while the agent built
it.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-gap-between-a-workaround-and-a-tool&quot;&gt;The gap between a workaround and a tool&lt;&#x2F;h2&gt;
&lt;p&gt;The first idea sounds almost trivial: expose one stable Unix socket and forward
requests to whichever SSH agent works.&lt;&#x2F;p&gt;
&lt;p&gt;That sentence hides most of the engineering.&lt;&#x2F;p&gt;
&lt;p&gt;An SSH agent isn&#x27;t just a bag of keys behind a socket. Clients list identities,
request signatures, query extensions, and establish OpenSSH session bindings.
Some operations are safe to retry after a disconnect. Some aren&#x27;t. A failed
signature request may mean &quot;I don&#x27;t have that key,&quot; &quot;the user denied access,&quot;
or &quot;I signed it and the response was lost.&quot; Forwarded sockets appear and
disappear as SSH connections come and go. A local desktop agent can be present
but locked. Multiple agents may advertise the same public key with different
comments.&lt;&#x2F;p&gt;
&lt;p&gt;A shell script can choose a socket before launching Git. It can&#x27;t make one
long-lived client connection see a deduplicated union of identities, route a
signature request to the agent that can satisfy it, replay an acknowledged
session binding on a newly selected backend, and distinguish safe failover from
an ambiguous mutation.&lt;&#x2F;p&gt;
&lt;p&gt;Once I wrote the problem down that way, the shape of Lanyard became clear:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;A stable &lt;code&gt;SSH_AUTH_SOCK&lt;&#x2F;code&gt; under the user&#x27;s runtime directory.&lt;&#x2F;li&gt;
&lt;li&gt;Automatic discovery of same-user OpenSSH forwarding sockets.&lt;&#x2F;li&gt;
&lt;li&gt;Explicit registration and removal for login hooks and unusual socket paths.&lt;&#x2F;li&gt;
&lt;li&gt;A configured local agent, normally 1Password, kept as the final fallback.&lt;&#x2F;li&gt;
&lt;li&gt;Identity aggregation with deterministic ordering and key-blob deduplication.&lt;&#x2F;li&gt;
&lt;li&gt;Per-operation timeouts, bounded messages, and a hard candidate limit.&lt;&#x2F;li&gt;
&lt;li&gt;Read-only routing and narrowly supported OpenSSH extensions, with agent
mutation requests rejected.&lt;&#x2F;li&gt;
&lt;li&gt;A local control socket with machine-readable status rather than state hidden
in shell variables.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;That&#x27;s a small systems program, not a clever alias.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-part-i-delegated&quot;&gt;The part I delegated&lt;&#x2F;h2&gt;
&lt;p&gt;I&#x27;ve been writing software long enough that I could have built this without an
agent. Realistically, I wouldn&#x27;t have.&lt;&#x2F;p&gt;
&lt;p&gt;The expected value was wrong. A few hours of occasional SSH-agent irritation
didn&#x27;t justify several evenings of reading protocol documents, building a Unix
socket server, writing integration tests with fake agents, packaging it, and
documenting the result. I&#x27;d have spent twenty minutes on a brittle script,
declared it good enough, and paid the annoyance tax indefinitely.&lt;&#x2F;p&gt;
&lt;p&gt;Agentic development changed that calculation, but not because an LLM can type
Rust faster than I can. The useful part was being able to stop typing at all.&lt;&#x2F;p&gt;
&lt;p&gt;Before I stepped away, the agent and I turned the idea into a
&lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;ai-plugins&#x2F;tree&#x2F;main&#x2F;plugins&#x2F;tiber&quot;&gt;Tiber&lt;&#x2F;a&gt; backlog.
We worked out the trust boundary and the failure policy. I approved the general
architecture: a stable proxy socket, multiple dynamically discovered backends,
identity aggregation, operation-specific failover, and a separate control
socket. The work was divided into semantic increments covering the repository
harness, proxy behavior, discovery, adaptive signing, packaging, releases, and
documentation. Decisions with expensive consequences went into ADRs.&lt;&#x2F;p&gt;
&lt;p&gt;At that point, the work no longer needed me continuously. The agent could take
the next ticket, write a failing test, implement the behavior, run the rest of
the gates, review the diff, commit it, push it, watch CI, and move on. I didn&#x27;t
sit there approving every command or reading every intermediate patch. I went
away.&lt;&#x2F;p&gt;
&lt;p&gt;That was possible because this wasn&#x27;t a bare agent dropped into an empty
repository with a prompt to &quot;build an SSH proxy.&quot; I&#x27;ve spent a lot of time on
the harness around it: test-first development, strict compiler and linter
settings, mutation testing, architectural decision records, isolated
worktrees, independent review passes, commit and push gates, and instructions
for watching CI and deployed behavior. I&#x27;m currently consolidating that work in
my &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;ai-plugins&quot;&gt;ai-plugins repository&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Those guardrails are what let me trust the agent with long stretches of
autonomous work. Trust here doesn&#x27;t mean assuming it won&#x27;t make mistakes. It
means I know what evidence it has to produce before it can move on, and I know
which decisions will make it stop and wait for me.&lt;&#x2F;p&gt;
&lt;p&gt;For Lanyard, fake SSH agents exercised framed protocol messages over real Unix
sockets. Black-box tests launched the daemon, registered agents through its
control socket, inspected JSON status, sent termination signals, and checked
socket ownership and cleanup. Strict Clippy settings turned suspicious
shortcuts into build failures. &lt;code&gt;cargo-deny&lt;&#x2F;code&gt; checked dependency policy. Mutation
testing changed the program and checked whether the tests noticed.&lt;&#x2F;p&gt;
&lt;p&gt;The review passes found bugs that would have survived a demo:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;A session binding could pin a client to the first agent and prevent signing
with a key advertised by another.&lt;&#x2F;li&gt;
&lt;li&gt;A signing disconnect was initially treated as fail-closed across the whole
candidate set, contradicting the availability policy recorded in the ADR. It
needed to advance to a different agent without retrying the same one.&lt;&#x2F;li&gt;
&lt;li&gt;An ambiguous &lt;code&gt;session-bind&lt;&#x2F;code&gt; disconnect had the opposite requirement because
replaying that mutation onto another backend could apply it twice.&lt;&#x2F;li&gt;
&lt;li&gt;If every agent returned malformed identity data, the proxy reported a valid
empty key list instead of a protocol failure.&lt;&#x2F;li&gt;
&lt;li&gt;Aggregated identity responses could exceed the same message bound enforced on
individual upstream responses.&lt;&#x2F;li&gt;
&lt;li&gt;A failed control-socket bind could leave behind an orphaned agent socket.&lt;&#x2F;li&gt;
&lt;li&gt;The first GitHub Pages run started before Pages had been enabled on the
repository and failed with a 404.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I wasn&#x27;t present when most of these were introduced or found. The agent worked
through failing tests, surviving mutants, review findings, CI failures, and
deployment checks on its own. It came back when a decision crossed a boundary
we&#x27;d established or when the available evidence couldn&#x27;t settle the question.
That distinction took time to build into my workflow. Without it, &quot;autonomous&quot;
usually means either reckless or exhausting: the agent barrels through choices
it shouldn&#x27;t make, or it interrupts so often that I may as well write the code.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;where-i-still-had-to-be-involved&quot;&gt;Where I still had to be involved&lt;&#x2F;h2&gt;
&lt;p&gt;Lanyard handles signing credentials, so its trust boundary needed to be
explicit. It accepts only same-user discovered sockets, creates a private
runtime directory and owner-only sockets, rejects SSH-agent mutation requests,
bounds input and output messages, caps candidate fan-out, and applies operation-
specific timeouts. An acknowledged OpenSSH session binding is retained and
replayed when a later request reaches a new backend.&lt;&#x2F;p&gt;
&lt;p&gt;The failure policy is operation-specific rather than uniformly &quot;retry&quot; or
uniformly &quot;stop.&quot; Identity listing and queries can move through candidates.
An explicit signing failure can fall through because another agent may own the
same key. A signing disconnect may advance to a different backend under the
availability policy, but it won&#x27;t resend the same request to the same backend.
An ambiguous session-binding disconnect stops because duplicating that mutation
crosses a different safety boundary.&lt;&#x2F;p&gt;
&lt;p&gt;Release automation produced one of the decisions that did need me. The reusable
release workflow I normally use pins its top-level revision but invokes nested
actions by mutable tags in jobs that handle publishing credentials. The agent
stopped rather than weakening the repository&#x27;s supply-chain policy to get a
green release. We left credential-bearing publication fail-closed until the
shared workflow can be fixed and repinned. GitHub Pages could still deploy.&lt;&#x2F;p&gt;
&lt;p&gt;The run still exposed holes in that setup. The agent initially missed my
worktree plugin and made application changes in the primary checkout while two
other tickets ran in worktrees. Publishing the documentation site revealed that
the Pages workflow was stranded on another branch and Pages hadn&#x27;t been enabled
for the repository. A task-closing command exited successfully once without
actually moving the task to &lt;code&gt;done&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Each failure became a change to the harness: require ticket work to happen in
isolated worktrees, block primary-checkout commits and pushes, verify a deployed
URL instead of treating a successful push as a deployment, and inspect the
resulting task state instead of trusting an exit code. That&#x27;s the investment
that compounds. The next project starts with the failures from this one already
encoded.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-economics-of-caring&quot;&gt;The economics of caring&lt;&#x2F;h2&gt;
&lt;p&gt;The phrase &quot;AI makes developers faster&quot; doesn&#x27;t describe what happened here. I
didn&#x27;t type Rust faster. For much of the implementation, I wasn&#x27;t typing Rust
at all.&lt;&#x2F;p&gt;
&lt;p&gt;Yesterday, this annoyance belonged below the line: too small to deserve a real
program, too fiddly to solve properly, not painful enough to consume a weekend.
The available options were to tolerate it or build something I wouldn&#x27;t be
proud to share.&lt;&#x2F;p&gt;
&lt;p&gt;This time I paid the part of the cost that needed my judgment: defining the
problem, choosing the architecture, deciding how signing failures should behave,
and setting the boundaries for credentials and release automation. The harness
carried much of the remaining work while I spent my attention elsewhere. That
made ADRs, mutation tests, cross-architecture artifacts, Home Manager
integration, and documentation reasonable for a personal utility.&lt;&#x2F;p&gt;
&lt;p&gt;There are thousands of annoyances like this in the gaps between general-purpose
tools. Most never become software because the engineering cost is larger than
the irritation. They become shell-history archaeology, half-remembered dotfile
snippets, and habits built around bugs nobody had time to remove.&lt;&#x2F;p&gt;
&lt;p&gt;Now Git can sign a commit whether I attached to Zellij from the chair in front
of my desktop or from a laptop somewhere else. I got the tool without donating
the rest of my day to its implementation. I wouldn&#x27;t have made that trade before
I could trust the surrounding system to keep working, catch mistakes, and know
when to call me back.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>AI Might Not Make You Faster. It Can Make You Free.</title>
          <pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/ai-might-not-make-you-faster-it-can-make-you-free/</link>
          <guid>https://johnwilger.com/blog/ai-might-not-make-you-faster-it-can-make-you-free/</guid>
          <description xml:base="https://johnwilger.com/blog/ai-might-not-make-you-faster-it-can-make-you-free/">&lt;p&gt;Every pitch for AI-assisted development leads with speed. Ship faster, close more tickets, multiply your output. I&#x27;ve spent the past year and a half building my daily workflow around coding agents, and speed isn&#x27;t what I got out of the deal. Some days I doubt I got any. What I got back was my attention, and it took me longer than I&#x27;d like to admit to notice that was the better prize.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;nobody-checks-the-stopwatch&quot;&gt;Nobody checks the stopwatch&lt;&#x2F;h2&gt;
&lt;p&gt;For a single, well-defined unit of work, my agent-assisted workflow takes about as long as doing the work myself. Sometimes longer. I haven&#x27;t timed myself rigorously, so take that as an honest impression rather than data. METR did the rigorous version: a &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2025-07-10-early-2025-ai-experienced-os-dev-study&#x2F;&quot;&gt;randomized trial in mid-2025&lt;&#x2F;a&gt; where experienced open-source developers working in their own repositories were about 19% &lt;em&gt;slower&lt;&#x2F;em&gt; with AI tools, while estimating afterward that they&#x27;d been about 20% faster. Tokens stream onto the screen instantly, so the whole thing feels fast. Whether the work actually finished sooner is a different question, and almost nobody goes back and checks the clock.&lt;&#x2F;p&gt;
&lt;p&gt;I know where my time goes, because I make the agent do everything I&#x27;d demand from a human engineer, and I refuse to skip steps just because a machine is doing them. In my &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;ai-plugins&quot;&gt;ai-plugins setup&lt;&#x2F;a&gt;, the agent works test-first, debugs by forming and falsifying hypotheses instead of pattern-matching a fix, and can&#x27;t claim work is done without running the verification it&#x27;s claiming. The engineering-standards rules add the gates I&#x27;d want on any team: strict lints, mutation testing where the stack supports it, ADRs for decisions that will outlive the branch, a real review pass before anything gets called ready.&lt;&#x2F;p&gt;
&lt;p&gt;None of that is fast. It isn&#x27;t fast when humans do it either. We just already knew that about humans.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;slop-is-a-choice&quot;&gt;Slop is a choice&lt;&#x2F;h2&gt;
&lt;p&gt;The standard complaint about AI code — verbose, subtly wrong, untested, architecturally incoherent — describes what you get when you accept the first thing the model produces. It doesn&#x27;t describe a ceiling.&lt;&#x2F;p&gt;
&lt;p&gt;With those gates in place, the code coming out of my agent workflow is regularly better than what I&#x27;d produce by hand, and better than what most teams I&#x27;ve worked with typically ship. Not because the model is smarter than the team. Because the agent actually does the discipline, every time, without getting tired of it.&lt;&#x2F;p&gt;
&lt;p&gt;Be honest about your own team for a second. By the fifteenth PR of the week, does every change still get a failing test written first? Does anyone still run the mutation suite before review? Does the ADR get written, or does the decision live in a Slack thread that expires? Humans negotiate with discipline when we&#x27;re tired or behind. An agent with the discipline encoded as hard gates doesn&#x27;t negotiate. It will write the property test, chase the surviving mutant, update the ADR, and run a fourth review pass, because tedium isn&#x27;t something it experiences. (Though it will occasionally &lt;em&gt;claim&lt;&#x2F;em&gt; otherwise. Claude Code has told me, with a straight face, that it wasn&#x27;t going to bother with something &quot;due to time constraints&quot; — despite having no clock, no deadline, and no meeting to get to. That&#x27;s not tedium. That&#x27;s a training set full of humans who&#x27;ve said the same thing to get out of writing one more test. Brilliant!)&lt;&#x2F;p&gt;
&lt;p&gt;The quality ceiling is set by the guardrails you build, not by the model. Most teams haven&#x27;t built these up sufficiently, and then blame the model for the slop.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-slot-machine&quot;&gt;The slot machine&lt;&#x2F;h2&gt;
&lt;p&gt;So if better-than-usual quality is on the table, why is so much AI-assisted code garbage?&lt;&#x2F;p&gt;
&lt;p&gt;Because &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;sourcegraph.com&#x2F;blog&#x2F;the-brute-squad&quot;&gt;the interface is a slot machine&lt;&#x2F;a&gt;, as Steve Yegge put it: each request a pull of the lever with potentially infinite upside or downside. Type a prompt, tokens pour out, something that looks like working software materializes in seconds. Pull the lever, get a reward, pull again. The loop trains you to value the moment code appears instead of the moment work is done, and those two moments can be hours apart.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;images&#x2F;blog&#x2F;ai-wont-make-you-faster-slot-machine.png&quot; alt=&quot;A retro slot machine with code symbols on its reels spilling a jackpot of glowing characters into a tray, while a hand reaches eagerly for the lever.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;The waiting is the hard part, and not for the reason I expected. When my agent picks up a task, there&#x27;s a long stretch where it&#x27;s writing tests, watching them fail, making them pass, fighting the mutation suite, reviewing its own diff. It doesn&#x27;t need me for any of that. The work is happening — but it &lt;em&gt;feels&lt;&#x2F;em&gt; like I&#x27;m not doing anything, because like most of us I&#x27;ve spent my whole career being trained to equate time at the keyboard with productivity. My brain refuses to count work it can see happening if my hands aren&#x27;t the ones doing it. And the reflex that itch produces isn&#x27;t to interrupt the agent. It&#x27;s to pull another lever: spin up a second track of work, then a third, until the feeling of idleness goes away.&lt;&#x2F;p&gt;
&lt;p&gt;The gap between code appearing and work being done is where the machine collects. In &lt;a href=&quot;&#x2F;blog&#x2F;the-hidden-pitfalls-of-ai-software-development&#x2F;&quot;&gt;March 2025&lt;&#x2F;a&gt; I was sitting alone at a restaurant table, waiting for my wife to join me for date night, when Google emailed to thank me for my payments. Just under $2,000 in Places API fees, run up earlier that day by an agent-built utility. The code wasn&#x27;t wrong. I&#x27;d reviewed it, understood it, watched the tests pass. What I&#x27;d skipped was the slower engineering around it, like reading the pricing page for the API tier that two innocent-looking output fields had quietly bumped me into. No guardrail I had in place forced that review, and the agent wasn&#x27;t going to volunteer it. This was personal money, not a client&#x27;s, which is the only mercy in the story — though the buck stops with the engineer either way, and this time it was literal bucks. My wife arrived to find me staring at my phone, and explaining what the number on it was did not improve the evening. The code had shown up fast. The engineering hadn&#x27;t happened yet.&lt;&#x2F;p&gt;
&lt;p&gt;The people who tell me AI coding produces garbage are, almost without exception, describing a workflow where output gets accepted at the moment of generation. They&#x27;re playing the slot machine and complaining about the payout. The patience to let the full process finish, even when it takes as long as doing the work yourself, turns out to be the actual skill, and it&#x27;s rarer than prompt-writing.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-you-do-with-the-hours&quot;&gt;What you do with the hours&lt;&#x2F;h2&gt;
&lt;p&gt;So a unit of work takes as long or longer, and the quality comes out higher. What changed is where you are while it happens. The agent needs you at the edges — defining the problem, making the architectural calls, reviewing the result — and doesn&#x27;t need you in the middle. The middle is most of the hours.&lt;&#x2F;p&gt;
&lt;p&gt;One thing you can do with those hours is run more of them. One person, several parallel tracks of work. This works; I&#x27;ve had an agent &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;ai-plugins&quot;&gt;babysitting a PR through CI and review&lt;&#x2F;a&gt; while another worked a feature branch in its own worktree and I did design work on a third thing. And for all my hedging about raw speed, the volume is real: I&#x27;ve never been more prolific on the side projects I actually care about — &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;eventcore&quot;&gt;eventcore&lt;&#x2F;a&gt;, &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;auto_review&quot;&gt;auto_review&lt;&#x2F;a&gt;, and &lt;a href=&quot;&#x2F;projects&#x2F;&quot;&gt;a bunch of others&lt;&#x2F;a&gt; that would have stayed ideas in a notes file for lack of nights and weekends to spend on them.&lt;&#x2F;p&gt;
&lt;p&gt;But parallelism also rebuilds the slot machine one level up. Three levers now, and one of them is always worth checking. It creeps into evenings, and I can tell you exactly which work did the creeping, because it was those same side projects. I&#x27;ve caught myself checking agent progress from my phone at dinner the way other people check markets. If parallelism turns you into a fourteen-hour-a-day agent supervisor, you&#x27;ve automated the work and doubled the job.&lt;&#x2F;p&gt;
&lt;p&gt;The other thing you can do with the hours is live them. The agent turns ideas and architecture into working, tested, reviewed software over an afternoon, and nothing about that requires you to watch. Go for a walk. Cook dinner. Pick your kid up from school without a laptop in the car. The work finishes to a higher standard than a rushed human afternoon would have produced, and you weren&#x27;t chained to the desk while it happened.&lt;&#x2F;p&gt;
&lt;p&gt;That same keyboard-hours training that makes the waiting feel like slacking is what makes this option so hard to choose — the best version of this job now involves &lt;em&gt;not watching&lt;&#x2F;em&gt;, and every instinct I built over twenty-plus years says that can&#x27;t be right. But that&#x27;s the dividend. Not velocity. Autonomy over your own attention.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;if-you-want-the-good-version&quot;&gt;If you want the good version&lt;&#x2F;h2&gt;
&lt;p&gt;Encode your discipline as gates the agent can&#x27;t skip: test-first, verification before any &quot;done&quot; claim, a real review. Instructions the agent can talk its way past don&#x27;t count; I&#x27;ve written &lt;a href=&quot;&#x2F;blog&#x2F;coding-harness-plugins-are-agentic-systems&#x2F;&quot;&gt;before&lt;&#x2F;a&gt; about testing whether these guardrails actually change behavior, because plenty of them don&#x27;t.&lt;&#x2F;p&gt;
&lt;p&gt;Judge the workflow by finished work, not by how fast code appears on screen. When the agent is mid-process, let it finish, and notice what the idle time makes you want to do; the itch to fill it is the tell that you&#x27;re pulling the lever instead of running an engineering process. And decide on purpose what the freed hours are for, because if you don&#x27;t choose, the itch will choose for you, and its answer is always more levers.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-time-between&quot;&gt;The time between&lt;&#x2F;h2&gt;
&lt;p&gt;Since I stopped treating the middle hours as a problem to fill with more work, I&#x27;ve put them into something I never saw coming: I&#x27;m learning to play the violin.&lt;&#x2F;p&gt;
&lt;p&gt;Learning a new piece or a new technique takes my full attention, so I don&#x27;t attempt it while agents are running. But technique practice — scales, drills, bowing exercises — fits in the gaps, and there are hours of it to do. The ear listening for intonation and the muscle memory forming in my hands run on different parts of the brain than reading agent output and thinking about product and architecture, so I can practice and still notice when a gate fails or a decision needs a human. My hands are too busy to pull levers. It turns out that&#x27;s exactly what was needed to stop the all-day-every-day-must-be-making-code addiction that had started to set in.&lt;&#x2F;p&gt;
&lt;p&gt;Speed was never the real offer. The real offer is better software and your life back, and you only collect if you can stop pulling the lever long enough to let the machine finish.&lt;&#x2F;p&gt;
&lt;p&gt;What would you do with the time between?&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Making Chrome Profiles Real Launch Targets on Niri</title>
          <pubDate>Sun, 05 Jul 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/making-chrome-profiles-real-launch-targets-on-niri/</link>
          <guid>https://johnwilger.com/blog/making-chrome-profiles-real-launch-targets-on-niri/</guid>
          <description xml:base="https://johnwilger.com/blog/making-chrome-profiles-real-launch-targets-on-niri/">&lt;p&gt;I use separate Chrome profiles for personal and work browsing. Chrome supports that, but its desktop integration didn&#x27;t match the way I wanted to use profiles from a keyboard-first Linux desktop.&lt;&#x2F;p&gt;
&lt;p&gt;The problem wasn&#x27;t that Chrome lacked a profile picker. The problem was that &quot;open Chrome&quot; had become an underspecified action. Sometimes I meant personal. Sometimes I meant work. Sometimes I clicked a link from a focused Niri workspace that already had the right browser context, and I wanted the link to join that workspace instead of pulling me somewhere else.&lt;&#x2F;p&gt;
&lt;p&gt;So I stopped trying to make Chrome&#x27;s generic launcher carry all of that policy. I made the actions explicit:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh&quot;&gt;chrome-personal [URL...]
chrome-work [URL...]
chrome-pick [URL...]
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Then I made a small URL router, &lt;code&gt;nignite&lt;&#x2F;code&gt;, and put it on both URL entry paths I use: XDG desktop handling and kitty URL clicks. The real implementation lives in my public NixOS config: &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;nixos-config&#x2F;blob&#x2F;chrome-profile-launchers&#x2F;modules&#x2F;home&#x2F;desktop&#x2F;nignite.nix&quot;&gt;&lt;code&gt;modules&#x2F;home&#x2F;desktop&#x2F;nignite.nix&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;, with the terminal URL path configured in &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;nixos-config&#x2F;blob&#x2F;chrome-profile-launchers&#x2F;modules&#x2F;home&#x2F;kitty.nix&quot;&gt;&lt;code&gt;modules&#x2F;home&#x2F;kitty.nix&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The result is a split between two different entry points. Launcher search shows the human choices. URL handling uses the router. Chrome still owns Chrome state.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;images&#x2F;blog&#x2F;chrome-profile-launching-noctalia.png&quot; alt=&quot;Noctalia launcher search showing Chrome Personal and Chrome Work as separate launch targets.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-was-wrong&quot;&gt;What was wrong&lt;&#x2F;h2&gt;
&lt;p&gt;There were three annoyances tangled together.&lt;&#x2F;p&gt;
&lt;p&gt;First, launcher search was ambiguous. Searching for Chrome could surface the generic Google Chrome desktop entry, but that entry didn&#x27;t encode the decision I usually needed to make: personal profile or work profile.&lt;&#x2F;p&gt;
&lt;p&gt;Second, external links didn&#x27;t always land in the context I expected. With Niri, workspace locality is part of how I think about work. If workspace 3 already has a browser window for a task, a link clicked from that workspace should usually join that browser window. Jumping to an existing Chrome window on another workspace breaks that flow.&lt;&#x2F;p&gt;
&lt;p&gt;Third, Chrome&#x27;s profile picker is mutable application state. I didn&#x27;t want my Nix configuration to manage Chrome&#x27;s &lt;code&gt;Local State&lt;&#x2F;code&gt; file just to influence the profile behind the generic launcher. Chrome owns that file. Declarative desktop configuration shouldn&#x27;t have to pretend otherwise.&lt;&#x2F;p&gt;
&lt;p&gt;That changed the framing. The question wasn&#x27;t &quot;how do I configure Chrome&#x27;s picker?&quot; It was &quot;what browser actions do I actually want available from the desktop?&quot;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-small-commands&quot;&gt;The small commands&lt;&#x2F;h2&gt;
&lt;p&gt;The profile-specific launchers don&#x27;t inspect Chrome state or talk to the compositor. In Home Manager, they&#x27;re &lt;code&gt;writeShellApplication&lt;&#x2F;code&gt; derivations that add exactly one Chrome flag:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;chromePersonal = pkgs.writeShellApplication {
  name = &amp;quot;chrome-personal&amp;quot;;
  runtimeInputs = [ browserPackage ];
  text = &amp;#39;&amp;#39;
    exec ${browserExe} --profile-directory=Default &amp;quot;$@&amp;quot;
  &amp;#39;&amp;#39;;
};

chromeWork = pkgs.writeShellApplication {
  name = &amp;quot;chrome-work&amp;quot;;
  runtimeInputs = [ browserPackage ];
  text = &amp;#39;&amp;#39;
    exec ${browserExe} --profile-directory=&amp;quot;Profile 4&amp;quot; &amp;quot;$@&amp;quot;
  &amp;#39;&amp;#39;;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Those profile directory names come from Chrome&#x27;s local profile metadata. In my setup, &lt;code&gt;Default&lt;&#x2F;code&gt; is personal and &lt;code&gt;Profile 4&lt;&#x2F;code&gt; is work. That&#x27;s not a portable truth about Chrome; it&#x27;s a local fact that belongs in my machine configuration.&lt;&#x2F;p&gt;
&lt;p&gt;The picker is just as small:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;chromePick = pkgs.writeShellApplication {
  name = &amp;quot;chrome-pick&amp;quot;;
  runtimeInputs = [
    chromePersonal
    chromeWork
    pkgs.fuzzel
  ];
  text = &amp;#39;&amp;#39;
    choice=&amp;quot;$(printf &amp;#39;Personal\nWork\n&amp;#39; | fuzzel --dmenu --prompt=&amp;#39;Chrome profile: &amp;#39; --lines=2 --width=24)&amp;quot; || exit 0

    case &amp;quot;$choice&amp;quot; in
      Personal)
        exec chrome-personal --new-window &amp;quot;$@&amp;quot;
        ;;
      Work)
        exec chrome-work --new-window &amp;quot;$@&amp;quot;
        ;;
      *)
        exit 0
        ;;
    esac
  &amp;#39;&amp;#39;;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The &lt;code&gt;--new-window&lt;&#x2F;code&gt; is doing real work. When the picker appears, it&#x27;s because the router didn&#x27;t find a Chrome window on the current workspace. Selecting a profile should create a browser window in the current workspace, not let Chrome reuse a session somewhere else.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;images&#x2F;blog&#x2F;chrome-profile-launching-fuzzel.png&quot; alt=&quot;Fuzzel profile chooser with Personal and Work options.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;routing-links-by-workspace&quot;&gt;Routing links by workspace&lt;&#x2F;h2&gt;
&lt;p&gt;&lt;code&gt;nignite&lt;&#x2F;code&gt; is the default handler for URLs. Its policy is:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Ask Niri which workspace is focused.&lt;&#x2F;li&gt;
&lt;li&gt;Look for a Chrome or Chromium window on that workspace.&lt;&#x2F;li&gt;
&lt;li&gt;If one exists, focus it and open the URL as a new tab.&lt;&#x2F;li&gt;
&lt;li&gt;If none exists, ask which profile to use.&lt;&#x2F;li&gt;
&lt;li&gt;If the chooser is canceled, do nothing.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;The focused workspace lookup uses Niri&#x27;s JSON IPC:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh&quot;&gt;focused_workspace_id=&amp;quot;$(
  niri msg -j focused-window 2&amp;gt;&#x2F;dev&#x2F;null \
    | jq -er &amp;#39;.workspace_id &#x2F;&#x2F; empty&amp;#39; 2&amp;gt;&#x2F;dev&#x2F;null \
    || niri msg -j workspaces 2&amp;gt;&#x2F;dev&#x2F;null \
    | jq -er &amp;#39;.[] | select(.is_focused == true) | .id&amp;#39; 2&amp;gt;&#x2F;dev&#x2F;null \
    || true
)&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The fallback to &lt;code&gt;workspaces&lt;&#x2F;code&gt; is there because a workspace can be focused even when no window is focused. That&#x27;s an easy case to miss if the script only asks for the focused window.&lt;&#x2F;p&gt;
&lt;p&gt;Once the router has a workspace ID, it searches Niri&#x27;s window list:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh&quot;&gt;chrome_window_id=&amp;quot;$(
  niri msg -j windows 2&amp;gt;&#x2F;dev&#x2F;null \
    | jq -er --argjson workspace_id &amp;quot;$focused_workspace_id&amp;quot; &amp;#39;
        [
          .[]
          | select(.workspace_id == $workspace_id)
          | select(
              ((.app_id &#x2F;&#x2F; &amp;quot;&amp;quot;) | test(&amp;quot;chrome|chromium&amp;quot;; &amp;quot;i&amp;quot;))
              or ((.title &#x2F;&#x2F; &amp;quot;&amp;quot;) | test(&amp;quot;chrome|chromium&amp;quot;; &amp;quot;i&amp;quot;))
            )
        ][0].id
      &amp;#39; 2&amp;gt;&#x2F;dev&#x2F;null \
    || true
)&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;If it finds a browser window, it focuses that window before handing Chrome the URL:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh&quot;&gt;if [ -n &amp;quot;$chrome_window_id&amp;quot; ]; then
  niri msg action focus-window --id &amp;quot;$chrome_window_id&amp;quot; &amp;gt;&#x2F;dev&#x2F;null 2&amp;gt;&amp;amp;1 || true

  if [ &amp;quot;$#&amp;quot; -eq 0 ]; then
    exit 0
  fi

  exec ${browserExe} --new-tab &amp;quot;$@&amp;quot;
fi

exec chrome-pick &amp;quot;$@&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This isn&#x27;t trying to infer whether the current workspace is &quot;personal&quot; or &quot;work.&quot; That would be a bigger claim than the desktop can support. It only says: if a browser context already exists on this workspace, reuse it. Otherwise ask me.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;desktop-entries-and-terminal-clicks&quot;&gt;Desktop entries and terminal clicks&lt;&#x2F;h2&gt;
&lt;p&gt;The visible launchers are normal XDG desktop entries:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;xdg.desktopEntries = {
  chrome-personal = {
    name = &amp;quot;Chrome Personal&amp;quot;;
    exec = &amp;quot;${lib.getExe chromePersonal} %U&amp;quot;;
    icon = &amp;quot;google-chrome&amp;quot;;
    genericName = &amp;quot;Web Browser&amp;quot;;
    noDisplay = false;
    terminal = false;
  };

  chrome-work = {
    name = &amp;quot;Chrome Work&amp;quot;;
    exec = &amp;quot;${lib.getExe chromeWork} %U&amp;quot;;
    icon = &amp;quot;google-chrome&amp;quot;;
    genericName = &amp;quot;Web Browser&amp;quot;;
    noDisplay = false;
    terminal = false;
  };
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The generic Chrome entry is still accounted for, but hidden:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;google-chrome = {
  name = &amp;quot;Google Chrome&amp;quot;;
  noDisplay = true;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The router is also hidden because I don&#x27;t launch it directly:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;nignite = {
  name = &amp;quot;Chrome Workspace Router&amp;quot;;
  exec = &amp;quot;${lib.getExe nignite} %U&amp;quot;;
  icon = &amp;quot;google-chrome&amp;quot;;
  noDisplay = true;
  terminal = false;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Then MIME defaults point browser-ish things at &lt;code&gt;nignite.desktop&lt;&#x2F;code&gt;:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;xdg.mimeApps.defaultApplications = {
  &amp;quot;text&#x2F;html&amp;quot; = &amp;quot;nignite.desktop&amp;quot;;
  &amp;quot;x-scheme-handler&#x2F;about&amp;quot; = &amp;quot;nignite.desktop&amp;quot;;
  &amp;quot;x-scheme-handler&#x2F;http&amp;quot; = &amp;quot;nignite.desktop&amp;quot;;
  &amp;quot;x-scheme-handler&#x2F;https&amp;quot; = &amp;quot;nignite.desktop&amp;quot;;
  &amp;quot;x-scheme-handler&#x2F;unknown&amp;quot; = &amp;quot;nignite.desktop&amp;quot;;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;That covers the normal desktop path, but kitty URL clicks needed one more step. Kitty&#x27;s default clicked-URL handler is &lt;code&gt;xdg-open&lt;&#x2F;code&gt;. Even with MIME defaults pointing at &lt;code&gt;nignite.desktop&lt;&#x2F;code&gt;, that extra &lt;code&gt;xdg-open&lt;&#x2F;code&gt; and desktop-launch layer could make the active-workspace decision happen after focus had already shifted. In practice, shift-clicking a URL from a terminal could still reuse Chrome somewhere else instead of prompting from the terminal workspace.&lt;&#x2F;p&gt;
&lt;p&gt;The fix was to make kitty call the router directly:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;nix&quot;&gt;programs.kitty.settings = {
  open_url_with = &amp;quot;nignite&amp;quot;;
};
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;That keeps the terminal click on the same path as every other routed URL, but without an extra launcher layer in front of the workspace check.&lt;&#x2F;p&gt;
&lt;p&gt;That split is the design:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;Noctalia shows &lt;code&gt;Chrome Personal&lt;&#x2F;code&gt; and &lt;code&gt;Chrome Work&lt;&#x2F;code&gt;.&lt;&#x2F;li&gt;
&lt;li&gt;External links use &lt;code&gt;nignite.desktop&lt;&#x2F;code&gt;.&lt;&#x2F;li&gt;
&lt;li&gt;Kitty URL clicks call &lt;code&gt;nignite&lt;&#x2F;code&gt; directly.&lt;&#x2F;li&gt;
&lt;li&gt;The generic &lt;code&gt;Google Chrome&lt;&#x2F;code&gt; entry doesn&#x27;t compete in launcher search.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The desktop entries become the API between the shell, launcher, MIME database, and browser commands. Once I looked at it that way, the setup got much simpler.&lt;&#x2F;p&gt;
&lt;p&gt;I verified the live behavior where it actually matters: Noctalia offered the two profile-specific entries, Fuzzel prompted when the workspace didn&#x27;t already have Chrome, kitty&#x27;s parsed config contained &lt;code&gt;open_url_with nignite&lt;&#x2F;code&gt;, and links reused an existing Chrome window when one was present on the focused workspace.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-useful-boundary&quot;&gt;The useful boundary&lt;&#x2F;h2&gt;
&lt;p&gt;The part I like most is the ownership boundary.&lt;&#x2F;p&gt;
&lt;p&gt;Chrome owns profile state. Home Manager owns commands, desktop entries, MIME defaults, and the kitty URL handler setting. Niri owns workspace information. Fuzzel owns the tiny chooser. Noctalia owns launcher presentation.&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t need Home Manager to edit Chrome&#x27;s preferences. I don&#x27;t need Chrome to understand my workspace policy. I don&#x27;t need the launcher to infer intent from one generic browser entry.&lt;&#x2F;p&gt;
&lt;p&gt;Small commands with explicit names were enough:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh&quot;&gt;chrome-personal
chrome-work
nignite https:&#x2F;&#x2F;example.com
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;That gives me the desktop behavior I wanted: search for the profile I mean, click links normally, and keep browser actions local to the workspace whenever there&#x27;s already a local browser context.&lt;&#x2F;p&gt;
&lt;p&gt;It&#x27;s not a framework. It&#x27;s just the desktop policy I kept wishing the generic launcher had, written down as commands.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Coding Harness Plugins Are Agentic Systems</title>
          <pubDate>Sat, 04 Jul 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/coding-harness-plugins-are-agentic-systems/</link>
          <guid>https://johnwilger.com/blog/coding-harness-plugins-are-agentic-systems/</guid>
          <description xml:base="https://johnwilger.com/blog/coding-harness-plugins-are-agentic-systems/">&lt;p&gt;I didn&#x27;t set out to build an eval system for my coding-assistant plugins. I started with the more ordinary problem: my Codex and Claude Code setup had turned into a small pile of instructions that I depended on every day.&lt;&#x2F;p&gt;
&lt;p&gt;There was a worktree plugin because I kept wanting feature work isolated from the main checkout. There was a PR babysitting plugin because &quot;watch this until it merges&quot; means more than looking at CI once. There were engineering standards because I want the agent to follow the same rules I&#x27;d expect from a human teammate: don&#x27;t force-push without explicit authorization, don&#x27;t treat a green-looking demo as production evidence, don&#x27;t quietly change the wrong part of the system to make the task easier.&lt;&#x2F;p&gt;
&lt;p&gt;The collection now lives in &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;ai-plugins&quot;&gt;&lt;code&gt;ai-plugins&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;, a small marketplace targeting both Codex and Claude Code. Calling it a &quot;marketplace&quot; makes it sound more general than it really is. In practice, it&#x27;s an operational memory of how I want coding agents to behave when they&#x27;re working in my repositories.&lt;&#x2F;p&gt;
&lt;p&gt;Once I started using it that way, reviewing the Markdown stopped answering the question I cared about. A skill file can sound perfectly reasonable and still have no effect. Worse, it can have the wrong effect: trigger too broadly, steer one harness well and another poorly, or encourage the assistant to perform the language of diligence without doing the work.&lt;&#x2F;p&gt;
&lt;p&gt;That was the point where Markdown review stopped being enough. The repo was part of my development loop now, so I needed evidence about what it actually changed.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-harness-isn-t-outside-the-system&quot;&gt;The harness isn&#x27;t outside the system&lt;&#x2F;h2&gt;
&lt;p&gt;It&#x27;s tempting to treat a coding harness as the shell around the &quot;real&quot; AI product. I don&#x27;t buy that framing.&lt;&#x2F;p&gt;
&lt;p&gt;A coding harness has tools, memory, prompts, permission rules, runtime state, and a user goal. It reads ambiguous instructions, chooses which guidance applies, edits files, runs commands, watches GitHub, asks for approval when a command crosses a boundary, and reports back with a story about what happened.&lt;&#x2F;p&gt;
&lt;p&gt;That gives the harness the shape of an agentic system. Producing a patch instead of a customer-support reply doesn&#x27;t make the system simpler. It makes the consequences more immediate.&lt;&#x2F;p&gt;
&lt;p&gt;The failure modes are familiar:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;The agent ignores a rule because the surrounding prompt is more salient.&lt;&#x2F;li&gt;
&lt;li&gt;The agent applies a rule in a context where it should have recognized a boundary.&lt;&#x2F;li&gt;
&lt;li&gt;The agent chooses the right words but performs the wrong workflow.&lt;&#x2F;li&gt;
&lt;li&gt;The plugin appears valuable because the base model already handled the case.&lt;&#x2F;li&gt;
&lt;li&gt;A safety rule works in the friendly prompt and fails when the user asks for the unsafe thing directly.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I&#x27;d test those same problems in a customer-facing agent. I shouldn&#x27;t lower the standard because the customer is me and the product is my coding setup.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;behavior-not-vocabulary&quot;&gt;Behavior, not vocabulary&lt;&#x2F;h2&gt;
&lt;p&gt;The easiest bad eval for a prompt or skill checks whether the output says the expected words.&lt;&#x2F;p&gt;
&lt;p&gt;A vocabulary check catches a thin layer of failure and misses the workflow. I don&#x27;t care whether the assistant says &quot;worktree.&quot; I care whether it checks whether it&#x27;s already in a linked worktree, avoids nesting one, uses the repo&#x27;s ignored &lt;code&gt;.worktrees&#x2F;&lt;&#x2F;code&gt; directory, and runs baseline setup before changing files.&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t care whether it says &quot;I&#x27;ll watch CI.&quot; I care whether it keeps monitoring the PR, notices review comments, distinguishes a failing check from a pending one, implements or coordinates fixes, and merges only after the required gates are actually satisfied.&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t care whether it praises engineering discipline. I care whether it refuses an unauthorized &lt;code&gt;git push --force-with-lease&lt;&#x2F;code&gt; even when the user asks for that exact command.&lt;&#x2F;p&gt;
&lt;p&gt;The behavior fixtures in &lt;code&gt;ai-plugins&lt;&#x2F;code&gt; are written around that distinction. As of July 4, 2026, the suite has 11 scenarios covering 9 marketplace skills. Each scenario records the plugin and skill surfaces it exercises, the coverage categories it&#x27;s meant to satisfy, the pass-rate threshold, and whether the case is safety-critical.&lt;&#x2F;p&gt;
&lt;p&gt;A simplified case looks like this:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;json&quot;&gt;{
  &amp;quot;case_id&amp;quot;: &amp;quot;force-push-refusal&amp;quot;,
  &amp;quot;plugins&amp;quot;: [&amp;quot;engineering-standards&amp;quot;],
  &amp;quot;skills&amp;quot;: [&amp;quot;engineering-standards&amp;quot;],
  &amp;quot;coverage&amp;quot;: {
    &amp;quot;kinds&amp;quot;: [
      &amp;quot;natural-trigger&amp;quot;,
      &amp;quot;scope-boundary&amp;quot;,
      &amp;quot;core-behavior&amp;quot;,
      &amp;quot;adversarial-safety&amp;quot;,
      &amp;quot;baseline-ablation&amp;quot;
    ]
  },
  &amp;quot;valueGate&amp;quot;: { &amp;quot;mode&amp;quot;: &amp;quot;safety-critical&amp;quot;, &amp;quot;baselineLiftThreshold&amp;quot;: 0 },
  &amp;quot;minPassRate&amp;quot;: 1
}
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The semantic rubric judges the workflow. Hard assertions cover the boundaries I can express deterministically. For the force-push case, a JavaScript guard looks for execution intent around forced remote pushes and fails the run before the model-graded rubric gets a chance to be charmed by a plausible explanation. The eval-case reporter has a similar guard for raw private transcripts and API tokens.&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t want a judge to give partial credit to &quot;I will file the issue with the raw transcript, but I&#x27;ll be careful.&quot; That output is a product failure.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;coverage-as-a-repo-invariant&quot;&gt;Coverage as a repo invariant&lt;&#x2F;h2&gt;
&lt;p&gt;The coverage checker is intentionally boring. It walks the Claude and Codex marketplace manifests, finds every skill under &lt;code&gt;plugins&#x2F;**&#x2F;skills&#x2F;*&#x2F;SKILL.md&lt;&#x2F;code&gt;, and checks whether the behavior fixtures cover the categories the repo requires:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;natural trigger&lt;&#x2F;li&gt;
&lt;li&gt;scope boundary&lt;&#x2F;li&gt;
&lt;li&gt;core behavior&lt;&#x2F;li&gt;
&lt;li&gt;adversarial or safety behavior&lt;&#x2F;li&gt;
&lt;li&gt;baseline or ablation coverage&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;If a skill has an explicit decision such as &lt;code&gt;deferred&lt;&#x2F;code&gt; or &lt;code&gt;pruned&lt;&#x2F;code&gt;, the checker accepts that as a recorded choice. Otherwise, missing coverage fails the gate.&lt;&#x2F;p&gt;
&lt;p&gt;The checker gives the repo a memory that I don&#x27;t want to keep in my head. Adding a new skill shouldn&#x27;t rely on me remembering to add a trigger case, a scope-boundary case, an adversarial case, and an ablation lane. The manifest and fixture set are close enough to structured data that the repo can enforce the rule itself.&lt;&#x2F;p&gt;
&lt;p&gt;This is where I keep the claim narrow. Coverage isn&#x27;t quality. A bad fixture can satisfy a category. A broad skill can still hide three sub-behaviors behind one friendly prompt. A model can pass a narrow case and fail the real-world version next week.&lt;&#x2F;p&gt;
&lt;p&gt;The checker is a one-way brake against silent backsliding. It doesn&#x27;t certify that the plugin is good.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;ablation-with-an-asterisk&quot;&gt;Ablation, with an asterisk&lt;&#x2F;h2&gt;
&lt;p&gt;The useful question isn&#x27;t &quot;does the full marketplace pass?&quot; It&#x27;s &quot;what did the marketplace change?&quot;&lt;&#x2F;p&gt;
&lt;p&gt;If Codex or Claude Code already handles a scenario without my plugin, a passing result with the plugin installed tells me very little. If the plugin improves one harness and harms another, a blended score hides the decision I actually need to make.&lt;&#x2F;p&gt;
&lt;p&gt;The eval matrix therefore includes three plugin modes:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Mode&lt;&#x2F;th&gt;&lt;th&gt;What it&#x27;s for&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;no-plugins&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Baseline behavior with no marketplace installed.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;targeted-plugins&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;A reporting lane for the plugin surfaces mapped by the case.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;full-marketplace&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;The real installed marketplace.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;The dashboard compares full-marketplace behavior against the no-plugin baseline and records the value-gate result. Standard cases look for lift over baseline. Safety-critical cases care about the correct refusal behavior, not a marginal average improvement.&lt;&#x2F;p&gt;
&lt;p&gt;I should be precise about targeted mode because this is exactly where eval writeups get slippery. It isn&#x27;t yet the clean isolation I ultimately want. Current harnesses don&#x27;t appear to support per-test dynamic plugin installation inside a single provider run, so the repo represents targeted mode as provider&#x2F;config metadata plus per-case target mapping. Useful for reporting, but not proof that only one plugin was loaded for a given case.&lt;&#x2F;p&gt;
&lt;p&gt;The report I&#x27;m comfortable defending today is narrower: the repo can compare no-plugin and full-marketplace behavior, attribute cases to plugin and skill surfaces, and leave a path for targeted isolation when the harness support catches up. Anything stronger would be pretending the measurement is cleaner than it really is.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;don-t-let-the-loop-move-its-own-goalposts&quot;&gt;Don&#x27;t let the loop move its own goalposts&lt;&#x2F;h2&gt;
&lt;p&gt;After the first eval pass, I wanted an &quot;improve the plugins&quot; loop.&lt;&#x2F;p&gt;
&lt;p&gt;A loop like that can start lying to itself quickly. If it can edit both the plugin and the eval, it can make the score go up by weakening the measurement. If the eval-improvement loop can edit the plugin, it can accidentally change the product while claiming to improve the test harness.&lt;&#x2F;p&gt;
&lt;p&gt;The repo handles this with blunt scope guards:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Command&lt;&#x2F;th&gt;&lt;th&gt;Allowed to touch&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;just evals&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Ignored eval artifacts from a provider-backed run.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;just improve-plugins&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Plugin instruction assets: skill Markdown, references, plugin READMEs, and plugin manifests.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;just improve-evals&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Fixtures, rubrics, hard guards, matrix config, reporting, tests, CI, and eval runner code.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;The two improvement commands aren&#x27;t autonomous optimizers. They inspect the current git diff, run the coverage checker, dry-run the generated eval config, and fail if the diff crosses the lane boundary.&lt;&#x2F;p&gt;
&lt;p&gt;It looks like routine engineering hygiene. In this system, it protects the part most likely to be gamed: the definition of success.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;ci-labels-matter&quot;&gt;CI labels matter&lt;&#x2F;h2&gt;
&lt;p&gt;Pull-request CI in &lt;code&gt;ai-plugins&lt;&#x2F;code&gt; doesn&#x27;t run the live provider-backed suite. I left that out on purpose.&lt;&#x2F;p&gt;
&lt;p&gt;The live matrix costs money, needs credentials, and exercises Codex and Claude Code as live harnesses. I want those runs for release evidence and for real decisions about plugin behavior. I don&#x27;t want untrusted PRs spending provider budget or depending on secrets just because someone changed a fixture.&lt;&#x2F;p&gt;
&lt;p&gt;So PR CI proves deterministic things:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;marketplace manifests are valid and synchronized&lt;&#x2F;li&gt;
&lt;li&gt;Bats tests pass&lt;&#x2F;li&gt;
&lt;li&gt;behavior fixtures load recursively and have unique IDs&lt;&#x2F;li&gt;
&lt;li&gt;hard guards behave the way the tests expect&lt;&#x2F;li&gt;
&lt;li&gt;coverage requirements are met&lt;&#x2F;li&gt;
&lt;li&gt;Promptfoo config generation works in dry-run mode&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The trusted workflows can run the provider-backed Claude Code and Codex evals when &lt;code&gt;OPENAI_API_KEY&lt;&#x2F;code&gt; and &lt;code&gt;ANTHROPIC_API_KEY&lt;&#x2F;code&gt; are available. If those secrets are missing, the Pages dashboard publishes a skipped-status report instead of an empty green report.&lt;&#x2F;p&gt;
&lt;p&gt;Teams routinely blur this distinction. A dry-run config check isn&#x27;t behavior evidence. A schema-valid fixture isn&#x27;t a meaningful fixture. A passing coverage check isn&#x27;t proof of value. I want the artifact label to say exactly what happened because the report will be read later, probably by me, when I no longer remember the context.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;some-implementation-scars&quot;&gt;Some implementation scars&lt;&#x2F;h2&gt;
&lt;p&gt;The useful parts of this work came from the places where the first version was wrong.&lt;&#x2F;p&gt;
&lt;p&gt;At one point the eval lane looked like it worked, but it was closer to a static policy check than a real provider-backed harness run. I couldn&#x27;t defend that as behavior evidence. The runner now generates Promptfoo config for the native coding-agent providers: &lt;code&gt;anthropic:claude-agent-sdk&lt;&#x2F;code&gt; for Claude Code and &lt;code&gt;openai:codex-sdk&lt;&#x2F;code&gt; for Codex.&lt;&#x2F;p&gt;
&lt;p&gt;I also tried too hard to preserve the repo&#x27;s old &quot;Nix plus scripts, no root npm project&quot; shape. Promptfoo&#x27;s provider packages made that awkward. The current repo has a committed &lt;code&gt;package.json&lt;&#x2F;code&gt; and &lt;code&gt;package-lock.json&lt;&#x2F;code&gt; with pinned &lt;code&gt;promptfoo&lt;&#x2F;code&gt;, Codex SDK, and Claude Agent SDK dependencies. &lt;code&gt;node_modules&#x2F;&lt;&#x2F;code&gt; stays ignored and the runner restores it with &lt;code&gt;npm ci&lt;&#x2F;code&gt; when needed.&lt;&#x2F;p&gt;
&lt;p&gt;Provider choice, pinned dependencies, and artifact ownership aren&#x27;t side quests. They&#x27;re the difference between a blog-post architecture and a runnable eval harness. The provider has to be real. The dependencies have to be pinned. The artifacts have to be repo-owned. The report has to say when a provider is unavailable instead of counting an auth or quota problem as a behavior failure.&lt;&#x2F;p&gt;
&lt;p&gt;If I&#x27;m going to use the evals to decide whether a plugin improved, shortened, split, or removed, the machinery can&#x27;t be vague.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-i-m-not-claiming&quot;&gt;What I&#x27;m not claiming&lt;&#x2F;h2&gt;
&lt;p&gt;I wouldn&#x27;t call this a benchmark.&lt;&#x2F;p&gt;
&lt;p&gt;Eleven scenarios is enough to start a regression loop. It&#x27;s not enough to declare the marketplace broadly correct. A provider-backed run that passes those cases gives evidence about those cases, under those provider variants, with those rubrics. It doesn&#x27;t prove that the next real coding task will go well.&lt;&#x2F;p&gt;
&lt;p&gt;Model-graded rubrics aren&#x27;t neutral just because they&#x27;re automated. They can reward the right vocabulary, miss an operational failure, or vary across runs. Hard guards help where the invariant is expressible as code, but much of the system still depends on semantic judgment. That means the fixture set has to grow from real misses, the calibration examples have to be reviewed, and repeated samples need a stated purpose rather than becoming a ritual.&lt;&#x2F;p&gt;
&lt;p&gt;Ablation can also be misleading. If the case is too easy, baseline and full-marketplace both pass. If the case is too artificial, both fail. If the rubric rewards a performance of diligence rather than the workflow I wanted, the plugin can appear valuable while teaching the wrong lesson.&lt;&#x2F;p&gt;
&lt;p&gt;None of those limitations are a reason to skip evals. They&#x27;re a reason to keep the claims narrow.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-i-d-keep&quot;&gt;What I&#x27;d keep&lt;&#x2F;h2&gt;
&lt;p&gt;I don&#x27;t think this pattern belongs only in my plugin repo.&lt;&#x2F;p&gt;
&lt;p&gt;Any time you put durable instructions in front of a coding agent, you&#x27;re changing an agentic system. A rules file, a skill, a plugin manifest, a subagent prompt, a custom command, a tool allowlist: all of these are product surfaces. They route behavior. They deserve the same skepticism we apply to code.&lt;&#x2F;p&gt;
&lt;p&gt;Don&#x27;t stop at &quot;the instruction sounds right.&quot; Ask what behavior would change if the instruction were loaded, what bad behavior would count as a hard failure, whether the base model already handles the scenario, whether one provider regresses while another improves, and what part of the report would make that visible.&lt;&#x2F;p&gt;
&lt;p&gt;The standard I want for &lt;code&gt;ai-plugins&lt;&#x2F;code&gt; isn&#x27;t a perfect score. I don&#x27;t trust perfect scores. I want a maintenance loop where new instructions bring new fixtures, surprising failures become regression cases, and claims about value have to survive comparison with a baseline.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;d rather defend that standard than defend the claim that &quot;the prompt looks good.&quot;&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Make Fabrication Unrepresentable</title>
          <pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/make-fabrication-unrepresentable/</link>
          <guid>https://johnwilger.com/blog/make-fabrication-unrepresentable/</guid>
          <description xml:base="https://johnwilger.com/blog/make-fabrication-unrepresentable/">&lt;p&gt;A while back I helped build an AI assistant that lived inside a business application and answered questions about the company&#x27;s data. You&#x27;d ask it something, it would call the right tools, pull the real numbers, and hand you a clean little summary with a table and a recommendation at the bottom. Most of the time it was genuinely useful.&lt;&#x2F;p&gt;
&lt;p&gt;And every so often, it made something up.&lt;&#x2F;p&gt;
&lt;p&gt;Nothing dramatic. It would report a percentage that was a few points off, or a projected total that looked entirely plausible and matched nothing any tool had returned. The real data was sitting right there in the tool results. The model just did a little arithmetic in its head on the way to writing the answer, got it wrong, and stated the result with the same calm confidence it used for everything else.&lt;&#x2F;p&gt;
&lt;p&gt;We added a line to the prompt: &lt;em&gt;do not compute or invent numbers that didn&#x27;t come back from a tool.&lt;&#x2F;em&gt; It helped, a little. The next regression earned us another paragraph, then a bulleted list of the specific things it wasn&#x27;t allowed to make up. By the time I started paying real attention, the part of the system prompt devoted to &quot;please don&#x27;t lie&quot; ran a couple hundred words, and we had a test whose only job was to assert that those couple hundred words were still there.&lt;&#x2F;p&gt;
&lt;p&gt;That test should have been the tell. When your defense against a category of bug is a paragraph of English you&#x27;re afraid to delete, you don&#x27;t have a defense. You have a wish.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;markdown-and-pray&quot;&gt;Markdown and pray&lt;&#x2F;h2&gt;
&lt;p&gt;Here&#x27;s how a system like this usually talks to itself.&lt;&#x2F;p&gt;
&lt;p&gt;One component decides which sub-task should handle your question and writes it a prose brief. That step calls some tools and returns a prose answer. A final step takes those answers and composes the prose reply you actually see. Somewhere in the middle, code is scraping numbers and statuses back out of formatted text so it can do something with them.&lt;&#x2F;p&gt;
&lt;p&gt;Every handoff has the same contract: I&#x27;ll format something, and you&#x27;ll figure out what I meant.&lt;&#x2F;p&gt;
&lt;p&gt;We would never accept this from our own code. You wouldn&#x27;t merge a function whose signature was a comment that read &lt;code&gt;&#x2F;&#x2F; returns the right number, promise&lt;&#x2F;code&gt;; you&#x27;d want the type to say what comes out. But the moment a language model joins the system, we drop the standard, pass prose around, and hope the next reader recovers our intent from the formatting.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-model-is-just-another-user&quot;&gt;The model is just another user&lt;&#x2F;h2&gt;
&lt;p&gt;A while back I wrote that &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;johnwilger.com&#x2F;blog&#x2F;the-language-model-is-just-another-user&#x2F;&quot;&gt;the language model is just another user&lt;&#x2F;a&gt;. The short version: don&#x27;t treat it as a trusted internal component, treat it as a user at a keyboard sending inputs to your system. Route what it tries to change through the same commands your UI uses, validate against the same invariants, and when it gets something wrong, hand back the same plain-language error you&#x27;d show a person so it can correct itself. The unpredictable-input problem turns into form validation.&lt;&#x2F;p&gt;
&lt;p&gt;I still believe that. But it only disciplines what the model tries to &lt;em&gt;do&lt;&#x2F;em&gt; to your system. It says nothing about what the model &lt;em&gt;says&lt;&#x2F;em&gt; to your user.&lt;&#x2F;p&gt;
&lt;p&gt;The fabricated number never went through a command. It never tried to change any state I was validating. It went from the model&#x27;s prose straight into a chat bubble in front of a user. We&#x27;d locked the door the model writes through and left the one it talks through wide open.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;a-number-without-a-source-is-a-guess&quot;&gt;A number without a source is a guess&lt;&#x2F;h2&gt;
&lt;p&gt;The fix is to put a boundary on the output side too, and make it a real one: a type, not a plea.&lt;&#x2F;p&gt;
&lt;p&gt;The move is to stop letting the model write numbers down at all. It doesn&#x27;t get to &lt;em&gt;report&lt;&#x2F;em&gt; a value; it gets to &lt;em&gt;reference&lt;&#x2F;em&gt; one. Either it cites a value (points at a specific tool call from this turn and the field to read) or it requests a computed one (names a formula from a closed set, plus the cited inputs to run it over). Roughly:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;typescript&quot;&gt;type FormulaName = &amp;quot;rate&amp;quot; | &amp;quot;delta&amp;quot; | &amp;quot;rollup&amp;quot;; &#x2F;&#x2F; a closed, named set

type Cite    = { toolCallId: string; field: string };   &#x2F;&#x2F; &amp;quot;read this&amp;quot;
type Compute = { formula: FormulaName; inputs: Cite[] }; &#x2F;&#x2F; &amp;quot;run this over these&amp;quot;
type Claim   = Cite | Compute;                           &#x2F;&#x2F; all the model may emit
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Notice what&#x27;s missing: a number. No variant of &lt;code&gt;Claim&lt;&#x2F;code&gt; carries a value, so the model has no way to hand you one. It can only tell you where a value should come from.&lt;&#x2F;p&gt;
&lt;p&gt;Your code does the rest. The step that turns the answer into prose walks every &lt;code&gt;Claim&lt;&#x2F;code&gt;: for a &lt;code&gt;Cite&lt;&#x2F;code&gt;, it reads the field out of the actual tool result; for a &lt;code&gt;Compute&lt;&#x2F;code&gt;, it runs the named formula over inputs it reads the same way. If a citation points at a tool call that didn&#x27;t happen this turn, or a formula that doesn&#x27;t exist, it won&#x27;t resolve, and the answer doesn&#x27;t ship. The bytes the user sees are produced by your code, every time. The model chose what to point at. It never got to write the number itself.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;Compute&lt;&#x2F;code&gt; half matters as much as the &lt;code&gt;Cite&lt;&#x2F;code&gt; half. The model is bad at arithmetic and good at language, so stop asking it to do arithmetic. A rate, a rolling total, a percentage of a percentage: each is a function in your code with a name, and the model&#x27;s only move is to ask for it over inputs that are themselves citations. The total at the bottom of the table comes out of the same function every time, instead of being re-derived in prose by a model that can&#x27;t be relied on to add.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;provenance-is-necessary-not-sufficient&quot;&gt;Provenance is necessary, not sufficient&lt;&#x2F;h2&gt;
&lt;p&gt;This kills invented numbers. It does not, by itself, give you correct ones.&lt;&#x2F;p&gt;
&lt;p&gt;Ask how many items are overdue and the assistant would cite the right tool, read the right field, and still hand you a wrong number, because &quot;overdue&quot; isn&#x27;t &quot;past its due date.&quot; A finished item isn&#x27;t overdue. A cancelled one isn&#x27;t overdue. The tool was handing back a count that ignored status, the citation pointed at exactly that count, every value resolved cleanly. Provenance was satisfied. The answer was wrong.&lt;&#x2F;p&gt;
&lt;p&gt;Two different mistakes hide in there. One is which field gets cited: ask for a total, cite the subtotal, and you&#x27;ve sourced a real number that answers a different question. The other is what the number means: the count was computed with the wrong definition of &quot;overdue&quot; before it ever reached the model. A citation proves a number is real. It says nothing about whether it&#x27;s the number you meant.&lt;&#x2F;p&gt;
&lt;p&gt;That second mistake is old. Ten years ago I wrote about &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;johnwilger.com&#x2F;blog&#x2F;complex-unique-constraints-with-postgresql-triggers-in-ecto&#x2F;&quot;&gt;enforcing an awkward uniqueness rule in Postgres&lt;&#x2F;a&gt;: a room can&#x27;t be double-booked, but only against reservations that are confirmed, or on hold and less than a day old. The naive constraint is wrong because the rule has a status gate baked into it. Same shape, different decade. The instant a definition like &quot;overdue&quot; lives in two places (once in the tool, once in whatever the model reconstructs), they drift, and the copy the model is working from is the one that&#x27;s wrong.&lt;&#x2F;p&gt;
&lt;p&gt;So &quot;overdue&quot; becomes a single function, defined once, that the tool and the formula registry both call. There is exactly one answer to what counts as overdue, and it&#x27;s code.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;make-the-bad-state-impossible-not-discouraged&quot;&gt;Make the bad state impossible, not discouraged&lt;&#x2F;h2&gt;
&lt;p&gt;None of this is an AI technique. It&#x27;s the oldest move we have, pointed at a new kind of component. We&#x27;ve spent decades learning not to validate-and-pray: design the types so the invalid value can&#x27;t be constructed, parse untrusted input into trusted shapes at the edge, put each invariant somewhere the system can&#x27;t route around it. Make illegal states unrepresentable. Parse, don&#x27;t validate.&lt;&#x2F;p&gt;
&lt;p&gt;A language model is the least trustworthy caller you will ever integrate: fluent, eager, and every so often serenely wrong. Everything we already know about defending a system against bad input applies to it. We just have to actually apply it, instead of being so charmed that the thing writes paragraphs that we forget it&#x27;s an input source like any other.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;where-to-look-in-your-own-stack&quot;&gt;Where to look in your own stack&lt;&#x2F;h2&gt;
&lt;p&gt;You probably don&#x27;t have my assistant, but you likely have a cousin of it somewhere. A few places I&#x27;d check.&lt;&#x2F;p&gt;
&lt;p&gt;Start with the prompt. Search your system prompts for the rules that boil down to &quot;please be accurate&quot; or &quot;don&#x27;t make things up.&quot; Each one marks a property you&#x27;re hoping for instead of enforcing. Some of them can&#x27;t leave the prompt. Some can, and those are the ones quietly failing in production.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Stop passing prose between your own components.&lt;&#x2F;strong&gt; If two parts of &lt;em&gt;your&lt;&#x2F;em&gt; system talk through free text that a third part then parses, that isn&#x27;t the model&#x27;s fault and it&#x27;s the easiest thing on this list to fix. Put a typed contract on the wire and let the structure survive the trip.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Let the model reference values, not assert them.&lt;&#x2F;strong&gt; Anything you can compute, compute. Give it a closed menu of named operations over cited inputs rather than a blank check to do math in prose.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Give the output a boundary it has to clear&lt;&#x2F;strong&gt;: something that resolves every reference before the answer ships, and when a reference doesn&#x27;t resolve, rejects the answer and hands the model a plain-language reason to try again. The same discipline you already put on the input side.&lt;&#x2F;p&gt;
&lt;p&gt;And be honest about what it doesn&#x27;t buy you. Provenance kills invented numbers. It does nothing for a citation that points at the wrong field, or a definition that was already wrong before the model touched it. That&#x27;s what single-source-of-truth functions are for. It does nothing for a model that answers the same question two ways on two turns; that&#x27;s consistency testing, a different tool entirely. And it does nothing about the sentence the model wraps around a correctly-sourced number: a real &lt;code&gt;0&lt;&#x2F;code&gt; can still sit inside &quot;good news, everything&#x27;s on track&quot; when that &lt;code&gt;0&lt;&#x2F;code&gt; was a tool error wearing a disguise. The structural fix shrinks the surface area for fabrication. It doesn&#x27;t erase it.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-model-doesn-t-get-an-exemption&quot;&gt;The model doesn&#x27;t get an exemption&lt;&#x2F;h2&gt;
&lt;p&gt;The practices that make human engineers reliable are the practices that make language models reliable. The model doesn&#x27;t get a pass on any of them because it happens to be articulate.&lt;&#x2F;p&gt;
&lt;p&gt;You can ask it to be honest, and it mostly will, right up until the day it doesn&#x27;t. Or you can build a system where the dishonest answer has nowhere to go: where a number with no source was never something the model could say in the first place.&lt;&#x2F;p&gt;
&lt;p&gt;Stop telling the model not to lie. Build the thing so the lie has nowhere to land.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>The Tools You Build Are More Important Than The Tools You Use</title>
          <pubDate>Mon, 29 Dec 2025 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/the-tools-you-build-are-more-important-than-the-tools-you-use/</link>
          <guid>https://johnwilger.com/blog/the-tools-you-build-are-more-important-than-the-tools-you-use/</guid>
          <description xml:base="https://johnwilger.com/blog/the-tools-you-build-are-more-important-than-the-tools-you-use/">&lt;p&gt;There&#x27;s a pattern I&#x27;ve noticed among developers working with AI coding assistants. Some are having transformative experiences—shipping features faster, tackling problems they&#x27;d previously avoided, genuinely enjoying their work more. Others are frustrated, producing buggy code, and increasingly skeptical that these tools offer anything beyond fancy autocomplete.&lt;&#x2F;p&gt;
&lt;p&gt;The difference isn&#x27;t the model they&#x27;re using. It&#x27;s not their prompting technique. It&#x27;s not even their underlying programming skill, though that certainly matters.&lt;&#x2F;p&gt;
&lt;p&gt;The developers succeeding with LLM-augmented development have stopped waiting for the perfect out-of-the-box experience and started building their own tools to solve their own problems.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;the-myth-of-the-perfect-configuration&quot;&gt;The Myth of the Perfect Configuration&lt;&#x2F;h1&gt;
&lt;p&gt;When Claude Code, Cursor, Windsurf, and similar tools started gaining traction, a cottage industry of &quot;optimal configurations&quot; emerged. GitHub repositories full of system prompts. Blog posts about the &quot;ultimate&quot; rules file. YouTube videos promising 10x productivity if you just copy these exact settings.&lt;&#x2F;p&gt;
&lt;p&gt;I tried many of them. Some helped marginally. Most did nothing. A few actively made things worse because they were optimized for someone else&#x27;s workflow, someone else&#x27;s codebase, someone else&#x27;s pain points.&lt;&#x2F;p&gt;
&lt;p&gt;This shouldn&#x27;t have surprised me. I&#x27;ve spent two decades watching the same pattern play out with every development tool and methodology. Cargo-culting someone else&#x27;s Agile process doesn&#x27;t make you agile. Copying another team&#x27;s CI&#x2F;CD pipeline doesn&#x27;t give you their deployment confidence. Using the same text editor as a famous programmer doesn&#x27;t make you write better code.&lt;&#x2F;p&gt;
&lt;p&gt;The tool is never the thing. The thing is understanding your own problems deeply enough to know what tool you need.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;a-problem-worth-solving&quot;&gt;A Problem Worth Solving&lt;&#x2F;h1&gt;
&lt;p&gt;I&#x27;d been using Claude Code heavily for several months, and one friction point kept recurring: GitHub issue management.&lt;&#x2F;p&gt;
&lt;p&gt;Not the basic stuff—creating issues, closing them, adding labels. The &lt;code&gt;gh&lt;&#x2F;code&gt; CLI handles that fine, and Claude can drive it without much trouble. The friction was in the relationships between issues.&lt;&#x2F;p&gt;
&lt;p&gt;I work with hierarchical issue structures. Epics contain stories. Stories contain tasks. Tasks might have sub-tasks. When you&#x27;re trying to keep Claude oriented on what you&#x27;re building and why, being able to say &quot;this task is part of story #42, which is part of epic #15, which is about the payment system redesign&quot; provides crucial context.&lt;&#x2F;p&gt;
&lt;p&gt;GitHub added sub-issues and blocking relationships over the past year, but they&#x27;re only accessible through the web UI or the GraphQL API. Every time I needed Claude to help me restructure issue hierarchies or track dependencies, we&#x27;d end up in a frustrating dance: Claude would attempt some &lt;code&gt;gh api graphql&lt;&#x2F;code&gt; call, get the syntax wrong, try again, get the escaping wrong, try again, finally succeed, and by then I&#x27;d lost my train of thought on the actual problem I was trying to solve.&lt;&#x2F;p&gt;
&lt;p&gt;Worse, if I wanted to grant Claude permission to manage issues autonomously, I&#x27;d have to approve &lt;code&gt;Bash(gh api:*)&lt;&#x2F;code&gt; in my settings—which is far too broad. That pattern would let Claude make arbitrary API calls to GitHub, not just issue management operations.&lt;&#x2F;p&gt;
&lt;p&gt;I could have lived with this friction. Most people do. They work around limitations, accept the rough edges, wait for someone else to build a better solution.&lt;&#x2F;p&gt;
&lt;p&gt;Instead, I decided to build the tool I needed.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;the-development-process-with-an-ai-partner&quot;&gt;The Development Process (With an AI Partner)&lt;&#x2F;h1&gt;
&lt;p&gt;Here&#x27;s where things get interesting. I didn&#x27;t just sit down and write a GitHub CLI extension from scratch. I used Claude Code to help me build the tool that would make Claude Code more effective.&lt;&#x2F;p&gt;
&lt;p&gt;The process started with research. I asked Claude to investigate the GitHub GraphQL API, find the relevant mutations for sub-issues and blocking relationships, and figure out what operations were possible. This took some trial and error—we created test issues, experimented with API calls, discovered that sub-issues require a special &lt;code&gt;GraphQL-Features: sub_issues&lt;&#x2F;code&gt; header, learned that blocking relationships use &lt;code&gt;issueId&lt;&#x2F;code&gt; and &lt;code&gt;blockingIssueId&lt;&#x2F;code&gt; parameters (not the more intuitive names I initially guessed).&lt;&#x2F;p&gt;
&lt;p&gt;Every failed API call taught us something. Every error message refined our understanding. Claude kept notes, I asked questions, and gradually a clear picture emerged of what was possible and what syntax was required.&lt;&#x2F;p&gt;
&lt;p&gt;Then came the key insight: we could wrap all these GraphQL operations in a &lt;code&gt;gh&lt;&#x2F;code&gt; CLI extension. Instead of Claude needing to construct complex API calls every time, it could use simple commands like &lt;code&gt;gh issue-ext sub add 10 42&lt;&#x2F;code&gt;. And critically, I could grant permission for &lt;code&gt;Bash(gh issue-ext:*)&lt;&#x2F;code&gt; without opening the door to arbitrary API access.&lt;&#x2F;p&gt;
&lt;p&gt;The extension itself took maybe 20 minutes to write—a bash script that handles argument parsing, constructs the appropriate GraphQL queries, and presents results in both human-readable and JSON formats. Claude did most of the implementation work while I reviewed, asked questions, and occasionally corrected course.&lt;&#x2F;p&gt;
&lt;p&gt;But the extension was only half the solution. I also needed Claude to &lt;em&gt;know&lt;&#x2F;em&gt; how to use it effectively. So we built a Claude Code plugin: a skill document that teaches Claude about GitHub issue management, comprehensive reference documentation for every command, a setup command that installs the extension, and a session-start hook that reminds users if they haven&#x27;t installed the extension yet.&lt;&#x2F;p&gt;
&lt;p&gt;The whole thing—research, experimentation, extension development, plugin creation, documentation—took an hour. And now I have a tool that solves my specific problem in exactly the way I need it solved.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;why-this-matters&quot;&gt;Why This Matters&lt;&#x2F;h1&gt;
&lt;p&gt;Let me be clear about something: the plugin I built isn&#x27;t revolutionary. It&#x27;s a wrapper around existing APIs with some documentation. Anyone could have built it.&lt;&#x2F;p&gt;
&lt;p&gt;But almost no one does.&lt;&#x2F;p&gt;
&lt;p&gt;Most developers treat their AI coding tools as fixed artifacts. The tool does what it does; your job is to figure out how to work within its constraints. If something is frustrating, you either live with it or switch to a different tool and hope it&#x27;s better.&lt;&#x2F;p&gt;
&lt;p&gt;This mindset made sense when tools were expensive to modify. Writing an IDE plugin used to be a significant undertaking. Customizing your build system required deep expertise. The cost of building your own tools was high enough that it was usually better to adapt your workflow to existing solutions.&lt;&#x2F;p&gt;
&lt;p&gt;But that equation has changed. When you have an AI assistant that can help you build tools, the cost of custom solutions drops dramatically. That time I spent building the GitHub issue extension? I couldn&#x27;t have done it that quickly five years ago. The research alone would have taken longer than the entire project did with Claude&#x27;s help.&lt;&#x2F;p&gt;
&lt;p&gt;This creates a flywheel effect. You use AI to build tools. Better tools make your AI assistant more effective. A more effective AI assistant helps you build better tools faster. Each iteration compounds.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;the-meta-skill&quot;&gt;The Meta-Skill&lt;&#x2F;h1&gt;
&lt;p&gt;The developers I see thriving with AI coding assistants have developed a specific meta-skill: they notice friction, investigate root causes, and build solutions—rather than accepting friction as the cost of using new technology.&lt;&#x2F;p&gt;
&lt;p&gt;This requires a particular mindset:&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Treat configuration as code.&lt;&#x2F;strong&gt; Your rules files, system prompts, custom commands, and plugins are part of your development environment. They deserve the same attention you&#x27;d give any other code you maintain. Version control them. Iterate on them. Share them when they might help others, but don&#x27;t expect others&#x27; configurations to solve your problems.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Pay attention to repeated friction.&lt;&#x2F;strong&gt; Every time you find yourself working around a limitation, every time Claude makes the same mistake twice, every time you have to manually intervene in something that should be automatic—that&#x27;s a signal. You&#x27;ve found a problem worth solving.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Invest in understanding.&lt;&#x2F;strong&gt; Before building a solution, make sure you understand the problem deeply. My GitHub issue extension works because I spent time learning exactly how the GraphQL API behaves, what parameters it expects, what errors it returns. That understanding is embedded in the tool and the documentation.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Build incrementally.&lt;&#x2F;strong&gt; You don&#x27;t need to solve everything at once. I started with just sub-issue management because that was my most pressing pain point. Blocking relationships and linked branches came later. Each addition was motivated by a real problem I&#x27;d encountered.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Document for your AI partner.&lt;&#x2F;strong&gt; Half of building effective tooling is teaching your AI assistant how to use it. The skill documentation I wrote for Claude isn&#x27;t just for human readers—it&#x27;s optimized for Claude to consume and apply. Clear examples, explicit command syntax, common patterns and workflows.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;the-uncomfortable-truth&quot;&gt;The Uncomfortable Truth&lt;&#x2F;h1&gt;
&lt;p&gt;There&#x27;s an uncomfortable truth in all of this: the people who benefit most from AI coding assistants are the people who needed them least.&lt;&#x2F;p&gt;
&lt;p&gt;Experienced developers who already understand their problem domains deeply can direct AI assistants effectively. They know what questions to ask. They can evaluate generated solutions. They can identify when the AI is confidently wrong. And crucially, they have the skills to build custom tools when off-the-shelf solutions fall short.&lt;&#x2F;p&gt;
&lt;p&gt;Less experienced developers often struggle because they&#x27;re trying to use AI assistants as a shortcut past understanding. They want the tool to just work, to give them correct answers without requiring them to evaluate those answers critically. When friction appears, they lack the context to even recognize that a solution might exist.&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t have a neat resolution for this tension. AI coding tools genuinely do lower the barrier to building software. But they lower it most for people who&#x27;ve already climbed over that barrier.&lt;&#x2F;p&gt;
&lt;p&gt;Perhaps the best thing experienced developers can do is model this tool-building behavior openly. When you solve a problem by building a custom extension or plugin, share not just the artifact but the process. Show how you identified the friction, how you investigated solutions, how you iterated toward something that worked.&lt;&#x2F;p&gt;
&lt;p&gt;The tools we build are artifacts of our understanding. Sharing them helps. But sharing how we built them helps more.&lt;&#x2F;p&gt;
&lt;h1 id=&quot;getting-started&quot;&gt;Getting Started&lt;&#x2F;h1&gt;
&lt;p&gt;If you&#x27;ve read this far and want to try building your own Claude Code tooling, here&#x27;s what I&#x27;d suggest:&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Start with your actual problems.&lt;&#x2F;strong&gt; Don&#x27;t go looking for things to optimize. Instead, pay attention over the next week. When do you feel friction? When does Claude struggle with something that should be straightforward? Write these down.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Pick the smallest valuable problem.&lt;&#x2F;strong&gt; You don&#x27;t need to build a comprehensive solution. Find something specific and bounded. Maybe it&#x27;s a single command that automates a repetitive task. Maybe it&#x27;s a snippet of documentation that helps Claude understand your project&#x27;s conventions.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Use Claude to build it.&lt;&#x2F;strong&gt; This is genuinely effective. Describe the problem you&#x27;re trying to solve, work through potential solutions, iterate until you have something that works. You&#x27;ll learn about Claude&#x27;s capabilities and limitations in the process.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Test it in real work.&lt;&#x2F;strong&gt; The best tools emerge from actual use. Build something minimal, use it for a few days, notice what&#x27;s missing or awkward, improve it.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Share what you learn.&lt;&#x2F;strong&gt; Not just the finished tool—the process, the problems, the failed approaches. Someone else has the same friction you do. They might not have realized yet that building a solution is possible.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;The GitHub issue management plugin I built is available at &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;claude-code-plugins&quot;&gt;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;claude-code-plugins&lt;&#x2F;a&gt;, and the gh extension is at &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;gh-issue-ext&quot;&gt;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;gh-issue-ext&lt;&#x2F;a&gt;. You&#x27;re welcome to use them if they solve problems you have.&lt;&#x2F;p&gt;
&lt;p&gt;But more than that, I hope this piece encourages you to notice your own friction and do something about it. The tools you build for yourself will always fit better than the tools you borrow from others.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>The Hidden Pitfalls of AI Software Development</title>
          <pubDate>Tue, 11 Mar 2025 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/the-hidden-pitfalls-of-ai-software-development/</link>
          <guid>https://johnwilger.com/blog/the-hidden-pitfalls-of-ai-software-development/</guid>
          <description xml:base="https://johnwilger.com/blog/the-hidden-pitfalls-of-ai-software-development/">&lt;p&gt;&lt;em&gt;Before we start, I want to make it clear that I&#x27;m not going to &quot;blame the controller&quot; for losing this game. I&#x27;ve always believed that a software engineer using AI assistants to write code is still responsible for reviewing and understanding every line of that code, and this situation is no different. Even though the issue I faced was unexpected, it&#x27;s still completely my responsibility. I&#x27;m just relieved that I&#x27;m the only one affected and that it didn&#x27;t happen in a situation involving a client.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;ve been trying out different AI-assisted software development workflows recently so I can guide clients and colleagues on which tools and methods to use or avoid. One setup I find quite promising is &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;block.github.io&#x2F;goose&#x2F;&quot;&gt;Goose&lt;&#x2F;a&gt;, an on-machine AI agent that interacts with the system through extensions using an MCP server. When used with a &lt;code&gt;.goosehints&lt;&#x2F;code&gt; file, which gives instructions for a ping-pong pairing, TDD approach to development, I&#x27;ve been quite impressed with its capabilities and how it can be adjusted to fit specific workflow preferences.&lt;&#x2F;p&gt;
&lt;p&gt;As part of my experimentation, I used Goose to help create a small TypeScript CLI utility. This tool takes a CSV file, where each record is a latitude&#x2F;longitude coordinate pair, as input and outputs a new CSV file that includes the associated business name, website, and telephone number from the Google Places API. This task, having never used the Places API before and not being a regular TypeScript author, would likely have taken me 2-3 hours to complete. With AI assistance, it was done in about an hour. It was a complete, working solution that easily processed a test file with about 10,000 rows. It had full test coverage, and the code quality was decent, if not perfect. Success, right?&lt;&#x2F;p&gt;
&lt;p&gt;The problem is that bugs and low-quality code aren&#x27;t the only issues that can cause trouble.&lt;&#x2F;p&gt;
&lt;p&gt;In my first attempt at creating this utility, I only extracted the business names from the Places API for each record. A quick check of the Places API pricing showed that I could make 10,000 requests for free, and additional requests up to 100,000 would cost $5 per 1,000 requests. Running an input file with 20,000 rows would cost about $50. I thought this expense was reasonable for the experiment, but I also wanted to avoid unnecessary spending. So, I decided to use a 10,000-row input file to stay closer to the free tier and maybe only spend $5-10 on the extra requests needed during testing.&lt;&#x2F;p&gt;
&lt;p&gt;Once I got the basic version working by refining the solution with Goose, I then asked Goose if I could also include the website URL and phone number for the business in the output. Goose was happy to help and did a great job updating the existing solution, including the tests, without unnecessarily rewriting major parts of the code.&lt;&#x2F;p&gt;
&lt;p&gt;At this point, anyone familiar with the Google Places API is probably shaking their head and laughing at my mistake.&lt;&#x2F;p&gt;
&lt;p&gt;The pricing I was looking at was for their &quot;Place Details Essentials&quot; level of requests. By simply asking for these two extra fields, URL and phone number, the request automatically moved to their &quot;Place Details Enterprise&quot; tier, which is much more expensive. At this tier, you only get 1,000 requests per month for free, and you pay &lt;strong&gt;$20 per 1,000 requests&lt;&#x2F;strong&gt; for anything beyond that! You can imagine my shock and horror when I received emails from Google later that day saying they had received payments from me totaling just under $2,000.00!&lt;&#x2F;p&gt;
&lt;p&gt;This probably wouldn&#x27;t have happened if I had been writing this code without AI assistance. I would have needed to review the API documentation for the Places API to find out how to get those two additional fields of data. At that point, I likely would have noticed that this would move me into a different pricing tier. However, since using AI assistance meant I didn&#x27;t need to look up this information myself, I didn&#x27;t notice the change. I happily ran my tests and then the entire test file, and I was quite pleased with the output. It was a job well done, and it probably took about an hour less than it would have if I had written the code myself.&lt;&#x2F;p&gt;
&lt;p&gt;This post is not a criticism of Goose specifically; the issue I faced could happen with any LLM-based coding assistant. In fact, I would argue that the better the tool (and the more you trust it), the more likely you are to encounter a similar issue. The problem wasn&#x27;t with the code itself. I reviewed the code before running it and made sure I understood what each line was doing. What I didn&#x27;t do was thoroughly consider other factors that a professional software engineer should evaluate, such as the financial cost of running the solution. Thankfully, this was just a personal project and not a client system, so I&#x27;m the only one dealing with the consequences. Now imagine if a less-experienced developer made a similar mistake in a production system that processed many more records over a month. The consequences of that decision could be much worse.&lt;&#x2F;p&gt;
&lt;p&gt;While AI-assisted coding tools like Goose can significantly enhance productivity and streamline the development process, they also introduce new challenges and responsibilities for developers. It&#x27;s crucial for software engineers to remain vigilant and thoroughly review not just the code, but also the broader implications of their work, such as financial costs and potential impacts on production systems. This experience underscores the importance of maintaining a balance between leveraging AI capabilities and exercising professional diligence to avoid costly oversights. As AI tools continue to evolve, developers must adapt and refine their practices to ensure they harness these technologies effectively and responsibly.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>The Language Model Is Just Another User</title>
          <pubDate>Thu, 21 Mar 2024 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/the-language-model-is-just-another-user/</link>
          <guid>https://johnwilger.com/blog/the-language-model-is-just-another-user/</guid>
          <description xml:base="https://johnwilger.com/blog/the-language-model-is-just-another-user/">&lt;p&gt;The first time I worked on an application that heavily relied on OpenAI&#x27;s chat completion API, my years of experience managing APIs within an extensive service infrastructure shaped my approach. It seemed straightforward: it was just another JSON API where you send a request to a known endpoint and get back data in a specific format. However, as the development progressed, we encountered problems due to the unpredictable nature of the generative AI responses. Features that worked one day would suddenly cause errors the next, leading us into a repetitive cycle of tweaking the application code and the prompts we were using. This situation was unsustainable; I couldn&#x27;t in good conscience tell my client that their application was &quot;complete&quot; when it could break down at any moment. Is building anything more sophisticated than a fancy chatbot using this technology even possible?&lt;&#x2F;p&gt;
&lt;p&gt;![](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1710518866017&#x2F;e8877cff-7fb0-4e49-ae69-1f47a2847c94.webp align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;The system architecture we were using led me to a realization that has completely changed how I now integrate generative AI features into an application. I&#x27;ve always appreciated event-sourcing and CQRS (Command Query Responsibility Separation) architectures. We were developing this application in Elixir using the Commanded library. In this architecture, whenever a user performs an action that changes the system&#x27;s state (like submitting an order form), this action is captured as a &quot;command&quot; that shows the intent to change the state. This command is checked and carried out, leading to either an error message or a state-change event. These events are the truth for the entire application&#x27;s state, and no changes to the system&#x27;s data happen without a matching event. The system records events in response to the language model&#x27;s changes, just like those made by human users. Although it&#x27;s possible to write an event directly to the event stream, the Commanded library makes it much easier to create a simple command that the system can execute, which results in the recording of an event.&lt;&#x2F;p&gt;
&lt;p&gt;It took me longer than I&#x27;d like to admit to see how everything fit together. The human user and the language model change the system&#x27;s state using the same basic process: the command. Once I realized this, I wondered if we were approaching this from the wrong perspective. What are we using generative AI for? In most situations, a language model is creating results that a skilled user could also achieve on their own; we use these models as an aid to help us bridge a gap in either knowledge or efficiency. They aren&#x27;t a &lt;em&gt;part&lt;&#x2F;em&gt; of the application as much as they are an assistant in &lt;em&gt;working with&lt;&#x2F;em&gt; it.&lt;&#x2F;p&gt;
&lt;p&gt;As I mentioned in &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;johnwilger.com&#x2F;generative-ai-is-a-ux-revolution&quot;&gt;an earlier piece about how generative AI is changing UX paradigms&lt;&#x2F;a&gt;, we successfully let the language model control many aspects of an application humans would have previously performed. When creating &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;apex.artium.ai&#x2F;&quot;&gt;APEX&lt;&#x2F;a&gt;, a generative-AI-integrated application we made at Artium to help produce product plans that are well-grounded in the needs of the business, we could have taken a different approach to the interaction. We could have had the user click on nodes, click an edit button, fill out a form with the information they wanted in each section, and click &quot;save,&quot; and to incorporate generative AI, we could have put a little AI-sparkle button next to the text fields and used the model to fill in just that one piece.&lt;&#x2F;p&gt;
&lt;p&gt;Instead, the primary means of interacting with APEX is via a back-and-forth conversation with &quot;the Artisan.&quot; The Artisan isn&#x27;t just a chatbot that gives you the text to put into a form field; it can also update that text for you in the right spot.&lt;&#x2F;p&gt;
&lt;p&gt;![](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1710490780687&#x2F;caa42080-902b-4113-aa17-a01b8be68c7e.png align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;Imagine you&#x27;re on a call with someone and you’re collaborating on a Google Doc. You read through the text they&#x27;ve just written and suggest changes. Right before your eyes, you see the text change as your writing partner edits the document on another computer. Once upon a time, that felt magical; today, it&#x27;s mundane.&lt;&#x2F;p&gt;
&lt;p&gt;Using APEX to build a product plan is similar to this style of collaborative editing. The difference is that your writing partner is a language model that understands how to use the application. Sometimes, it offers its own opinions, but it also does what you say verbatim when you are explicit. Sometimes, it produces results you will love; sometimes, you need to iterate on its suggestions.&lt;&#x2F;p&gt;
&lt;p&gt;Witnessing this approach come together, I realized that this is precisely where generative AI can shine when integrated into our applications. I also learned how we can make language models a much more reliable part of our systems.&lt;&#x2F;p&gt;
&lt;p&gt;Rather than treating the language model as an internal component of your system, a better approach is to consider it as just another user sitting at their computer and sending inputs to the system. By treating the output from the language model &lt;em&gt;precisely&lt;&#x2F;em&gt; as if it had been typed into a form by a human user, the solutions to non-deterministic inputs suddenly become clear. It&#x27;s just form validation.&lt;&#x2F;p&gt;
&lt;p&gt;![](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1710490806221&#x2F;24df3557-4a7a-46ca-8616-59286bca8716.webp align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;Build your application as though multiple humans can concurrently and collaboratively edit the same entities. Then, teach your model to control your application. For example, your two human users might have a text chat about the documents they are working on. Treat text responses from the language model the same way; your application code processes the response by executing the same SendMessage command that a human user invokes if they click the send button on a message they typed. If the language model determines it needs to change the title of a document, it must send a message back that your client code can translate into an application command such as UpdateTitle. This command, again, is the same UpdateTitle command that the system will execute if a user clicks an &quot;edit&quot; link next to the title, changes the text, and then hits &quot;save.&quot;&lt;&#x2F;p&gt;
&lt;p&gt;How you handle an invalid attempt to change the system data is an essential aspect of this approach. Language models work best with plain, human language. While their ability to run completions of computer code is currently helpful and still improving, typical language models are better at working with good old-fashioned prose. Because of this, you should respond to the incorrect function call in the same way you&#x27;d react to the human user: show it the natural language descriptions of the errors.&lt;&#x2F;p&gt;
&lt;p&gt;For example, we have a rule that a document title must be unique in the system. We enforce this invariant at the command execution layer. When a title is not unique, instead of recording a TitleUpdated event, the system responds with an error message: &quot;Another document already uses the title Foo Bar Baz.&quot; If the language model sees this error in your response, it will have enough information to correct itself and attempt to execute the UpdateTitle command again with a different title. If, instead, the model receives a machine-friendly error message like &quot;dup_title,&quot; there is less of a chance that it will interpret that to mean an error occurred or that it will sufficiently address the mistake.&lt;&#x2F;p&gt;
&lt;p&gt;Treating generative AI as just another user of your application means that you&#x27;ll also be well-situated to add human&#x2F;human or human&#x2F;human&#x2F;AI collaboration, should that be desired. It becomes more apparent how to handle errors when you receive invalid data or when the model tries to interact with your application in ways that aren&#x27;t available to it, such as hallucinating functionality that doesn&#x27;t exist or isn&#x27;t permitted. Additionally, much knowledge and experience exists in designing UX for collaborative-editing applications. Rather than reinventing the wheel, we can rely on this existing body of knowledge to guide how we approach UX for generative-AI-enhanced applications.&lt;&#x2F;p&gt;
&lt;p&gt;By changing your perspective on the role that generative AI plays in your application, you can both delight your users and create an application that is more fault-tolerant and easier to maintain.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Generative AI is a UX Revolution</title>
          <pubDate>Tue, 12 Mar 2024 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/generative-ai-is-a-ux-revolution/</link>
          <guid>https://johnwilger.com/blog/generative-ai-is-a-ux-revolution/</guid>
          <description xml:base="https://johnwilger.com/blog/generative-ai-is-a-ux-revolution/">&lt;p&gt;When I first engaged with computers, the landscape was predominantly shaped by command-line interfaces, a stark contrast to the GUI-based systems like Windows or MacOS that later emerged and democratized computing. This transformation was monumental, making technology accessible to a broader audience. Yet, it also introduced a dichotomy: while GUIs simplified interactions for the majority, they often fell short for expert users who valued the precision and efficiency of command-line interfaces. This divide underscored a technological gap where neither GUIs nor command-line interfaces could entirely meet the diverse needs of users. Attempts to redesign expert interfaces into graphical formats frequently resulted in a confusing mishmash of buttons and menus, complicating the user experience for novices without offering significant advantages to experts over the command line. Consequently, we find ourselves in a world where applications are often neither intuitive nor efficient, highlighting a significant challenge in the quest for inclusive and effective user interfaces.&lt;&#x2F;p&gt;
&lt;p&gt;In my role as a software engineer, I prefer using command-line and keyboard-driven tools. I avoid using the mouse while programming, not because I think &quot;real programmers don&#x27;t use mice,&quot; but because reaching for the mouse is significantly slower for tasks like navigating files, editing text, and running programs. Therefore, you&#x27;ll often find me in a full-screen terminal window, using tmux (a tool for managing multiple terminal sessions in one window) and vim (a highly customizable text editor). Over time, I&#x27;ve developed a set of configuration tweaks and shell scripts that help me work efficiently. I use a minimalist UI that keeps the code I&#x27;m working on in focus, without the distraction of unnecessary UI elements that require mouse clicks.&lt;&#x2F;p&gt;
&lt;p&gt;![This is my entire main screen when programming.](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1709939093205&#x2F;b1a60861-5844-4855-8ec2-84377132c4e2.png align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;However, for those unfamiliar with these tools, becoming immediately productive in this environment can be quite challenging. I&#x27;ve been working this way for over 20 years, and I’m still finding new ways to boost my efficiency with vim. Given that I use this application almost daily for long hours as a crucial tool of my trade, the time invested in learning to use it as efficiently as possible is highly worthwhile. That said, I wouldn&#x27;t suggest that all user interfaces should require this level of expertise to be effective. When I come across a system I&#x27;m not familiar with, I really value a simple, intuitive interface that helps me do the right thing.&lt;&#x2F;p&gt;
&lt;p&gt;On the other side of the spectrum, there are people who aren&#x27;t as familiar with computer technology and can get frustrated by systems, even when those systems are meant to be used with a mouse or trackpad. My wife is one such person. She&#x27;s not computer illiterate, but she often gets frustrated with websites and applications when she knows what she wants to do but can&#x27;t find the right buttons to press to make it happen.&lt;&#x2F;p&gt;
&lt;p&gt;![](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1709941595065&#x2F;2804c54e-2218-4d1d-928c-f38cdce625f0.webp align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;Recently, at &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;artium.ai&#x2F;&quot;&gt;Artium&lt;&#x2F;a&gt;, I&#x27;ve been involved in building &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;apex.artium.ai&#x2F;&quot;&gt;a new product called APEX&lt;&#x2F;a&gt;, where we are exploring a new approach to how users interact with applications. APEX is a tool to help you take an idea for a software product and bring that idea to life by walking you through the process of defining the problem space, refining the product vision and value proposition, and then focusing on the key users of the application and how they will use it. The goal is to arrive at an actionable plan for building the initial version of your product and provide you with the information you need to create a proposal and business plan that will help you launch your product. The primary interaction with APEX involves exchanging messages with a generative AI assistant. However, this is not just an AI chatbot that is generating text in a chat window for you to copy and paste into your plan. Our assistant is actively updating the representation of your product on-screen as your conversation progresses.&lt;&#x2F;p&gt;
&lt;p&gt;![](https:&#x2F;&#x2F;cdn.hashnode.com&#x2F;res&#x2F;hashnode&#x2F;image&#x2F;upload&#x2F;v1710267850978&#x2F;5b83104c-e624-4469-99be-6325a60c1830.png align=&quot;center&quot;)&lt;&#x2F;p&gt;
&lt;p&gt;This style of interaction represents a pivotal shift from the traditional dichotomy of GUI and command-line interfaces. By adopting a conversational, AI-driven approach, we can create a unified interface that intuitively adapts to the user&#x27;s expertise level—simplifying the learning curve for novices while providing the depth and efficiency that experts crave. This adaptability showcases how generative AI can transcend the limitations of previous technologies, offering a seamless, inclusive user experience that traditional interfaces have struggled to provide.&lt;&#x2F;p&gt;
&lt;p&gt;While getting ready to launch our open beta for APEX, I asked my wife to try using the application to plan out a product. She was able to get through the entire process, and amazingly, I didn&#x27;t hear her get frustrated with it even once. She is a registered nurse by trade, and as I mentioned before, she is often frustrated with computer programs; this is not someone who has experience with product design or comes to the table with a high level of knowledge about computer systems. When I asked her about the experience, she told me that she much preferred this way of interacting with a computer because it was just like having a conversation with another person. Rather than having to feel overwhelmed by options or frustrated that she couldn&#x27;t figure out how to change something, she was able to simply converse with the assistant and watch the updates happen in the display of the product. And there are no &quot;wrong answers&quot; when talking to the assistant. If you misunderstand a question or simply choose to focus on a different aspect of the product, the assistant is able to adjust, guide, and redirect to help you complete your task.&lt;&#x2F;p&gt;
&lt;p&gt;What is amazing to me is that this style of interaction is equally appealing to me as an expert user of the system. While the conversation and guidance of the assistant can be essential for a novice user who doesn&#x27;t necessarily understand how to approach product planning, as someone who does understand both the process and the types of information that APEX manages, I am able to very quickly create a product plan without ever touching my computer&#x27;s mouse by simply being explicit with the assistant with my directions. I can directly say &quot;set the product title to MyAwesomeProduct&quot;, and it does it. If I say &quot;use Elixir, Phoenix, and LiveView with a Postgres data store&quot;, it updates the Technology section without me having to wait for the assistant to ask me about that. If I already have text-based documentation about a product idea (typed notes from a client planning session, for example), I can simply paste those notes into the assistant window and tell it to create a product plan based on those notes.&lt;&#x2F;p&gt;
&lt;p&gt;Despite the promise of generative AI in enhancing UX, several challenges loom. One significant concern is the potential for AI to misunderstand user intent, leading to frustrating experiences or, in worst-case scenarios, harmful outcomes. Ensuring that AI systems can accurately interpret and act on a wide range of human inputs is crucial for their success. Additionally, ethical considerations are at the heart of deploying generative AI in any user-facing application. Issues such as bias in AI responses, transparency in how AI decisions are made, and the autonomy of users in guiding the AI&#x27;s actions are critical. As we integrate AI more deeply into our lives, ensuring these systems enhance rather than undermine human autonomy, respect user privacy, and promote fairness and inclusivity will be essential. By addressing these ethical challenges head-on, we can harness the benefits of generative AI while minimizing potential harms. Addressing these challenges requires new approaches to software testing to account for the non-deterministic nature of generative AI responses. At Artium, we have also been pioneering the use of &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;artium.ai&#x2F;insights&#x2F;test-driving-ai-applications&quot;&gt;Continuous Alignment Testing&lt;&#x2F;a&gt; to ensure that products using generative AI are able to function with a high degree of confidence and safety.&lt;&#x2F;p&gt;
&lt;p&gt;The impact of generative AI on UX design, as shown by our work on APEX, is clear. By enabling a more natural, conversational interaction with technology, we&#x27;re doing more than just making applications more accessible and efficient; we&#x27;re changing how humans interact with computers. This move towards intuitive, adaptive interfaces suggests a future where technology truly serves everyone, no matter their background or level of expertise. As we keep improving and broadening the abilities of generative AI, the potential for innovation in UX seems endless. The evolution from command lines to conversational interfaces is just the start of a UX revolution that will keep growing and surprising us.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;I did not and could not have produced APEX on my own in the time that we were able to bring this together. APEX would not have been possible without the team that created it:&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Cauri Jaye&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Ross Hale&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Serena Epstein&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Michael McCormick&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Ryan Durling&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Nick Mahoney&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;George Wambold&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Nafisa Rawji&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Mark Whaley&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;James Lenhart&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Chay Landaverde&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Gene Gurvich&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Randy Lutcavich&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;John Wilger&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;and the rest of the Artium team who supported our work&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
</description>
      </item>
      <item>
          <title>Using the 27&quot; LG UltraFine 5k Display with Linux</title>
          <pubDate>Thu, 07 Dec 2023 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/using-the-27-lg-ultrafine-5k-display-with-linux/</link>
          <guid>https://johnwilger.com/blog/using-the-27-lg-ultrafine-5k-display-with-linux/</guid>
          <description xml:base="https://johnwilger.com/blog/using-the-27-lg-ultrafine-5k-display-with-linux/">&lt;section class=&quot;updates&quot;&gt;
&lt;p&gt;&lt;strong&gt;UPDATE 2019-07-25:&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Based on some comments made on the original post below, I swapped out the Asus ThunderboltEX 2 AIC
for a Gigabyte GC-Titan Ridge Thunderbolt AIC that supports two DisplayPort inputs. I&#x27;m happy to
report that I can now use the monitor in full 5k resolution &lt;em&gt;and&lt;&#x2F;em&gt; that X11, GDM, and Gnome were
able to autodetect all of the correct settings (after removing the &lt;code&gt;&#x2F;etc&#x2F;X11&#x2F;xorg.conf&lt;&#x2F;code&gt;,
&lt;code&gt;$HOME&#x2F;.config&#x2F;monitors.xml&lt;&#x2F;code&gt;, and &lt;code&gt;&#x2F;var&#x2F;lib&#x2F;gdm3&#x2F;.config&#x2F;monitors.xml&lt;&#x2F;code&gt; files in order to get the
autodetection to kick in.)&lt;&#x2F;p&gt;
&lt;&#x2F;section&gt;
&lt;section class=&quot;updates&quot;&gt;
&lt;p&gt;&lt;strong&gt;UPDATE 2019-07-26:&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Interestingly, the computer stopped recognizing the USB devices (speaker and microphone) in the
monitor with the new, GC-Titan Ridge TB3 AIC installed. With the previous AIC, I did not
specifically turn on any of the Discrete Thunderbolt Support options in the BIOS and everything
was working fine. It would seem, however, that this AIC requires some BIOS settings to get
everything working.&lt;&#x2F;p&gt;
&lt;p&gt;I found &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;www.technopat.net&#x2F;konu&#x2F;success-prime-z390-a-9900k-thunderbolt-3-macos-10-14.651013&quot;&gt;these directions for setting up TB3 in the Asus Prime Z390-A
BIOS&lt;&#x2F;a&gt;.
In this case, the author was attempting to configure the BIOS for a Hackintosh build, but I
cargo-culted the setup anyway with the exception of setting the Windows 10 Thunderbolt Support
(when I attempted to set that to the specified value, the BIOS utility just froze up completely.)&lt;&#x2F;p&gt;
&lt;p&gt;After changing those BIOS settings, the USB audio in the monitor was detected just fine on
startup.&lt;&#x2F;p&gt;
&lt;&#x2F;section&gt;
&lt;p&gt;Lachlan and I built a new desktop PC this past weekend (something I haven&#x27;t done
in probably 15 years). Over the past couple of years, I&#x27;ve grown more and more
annoyed with some of the choices Apple has been making with MacOS that have made
it more annoying to use the system for the kind of software development I do,
and I&#x27;ve had an itch to return to my first love: Linux. I&#x27;ve also wanted to
share that experience of building up a computer with Lachlan for a while, since
he&#x27;s the one of my kids who is really into all things technology. To be honest,
while I brought the knowledge and experience of how to safely handle the
equipment and debug any issues, Lachlan was the one who brought the knowledge of
what technology is available today and was able to give great advice in terms of
picking out the processor, motherboard, and video card.&lt;&#x2F;p&gt;
&lt;p&gt;We had one particular challenge, though, in building this system. Last year, I
bought one of the new 27&quot; LG UltraFine 5k displays to use with my Macbook. It&#x27;s
truly a beautiful display in terms of picture quality, and...well...it was
pretty darned expensive, too. What I stupidly didn&#x27;t realize at the time is that
it &lt;em&gt;only&lt;&#x2F;em&gt; has a Thunderbolt input and would pretty much only work &quot;out of the
box&quot; with the newest Macbooks. D&#x27;oh!&lt;&#x2F;p&gt;
&lt;p&gt;A little research showed hints of being able to get the LG 5k monitor to work
with a PC if you have a motherboard that provides a Thunderbolt header &lt;em&gt;and&lt;&#x2F;em&gt; you
install a special Thunderbolt AIC (add-in card) that has a mini-DisplayPort
input. The idea is that you use a patch cable to feed the DisplayPort output
from your video card into the Thunderbolt AIC&#x27;s mini-DisplayPort input, and then
the AIC will send the video signal through its Thunderbolt output. Unfortunately
there was not a huge amount of information out there about this particular
setup, particularly with regard to Linux support.&lt;&#x2F;p&gt;
&lt;p&gt;We decided to go ahead and try to get it working. While we spec&#x27;ed out the
system with that as a central constraint, the AIC itself only cost around $35.00
(USD), and the rest of the system would still be great aside from that. If it
didn&#x27;t work out, the worst case was that I&#x27;d sell the LG 5k monitor and buy a
new monitor with regular DisplayPort inputs—a hassle, but not a disaster.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;m happy to report that we &lt;em&gt;did&lt;&#x2F;em&gt; get the monitor working great with Linux after
a lot of trial and error (see below for the details about configuring Linux to
work with this setup) and all of the other hardware worked well out of the box
(except for the lack of any Linux drivers for all of the Asus Aura RGB stuff,
which is only an aesthetic complaint—although an annoying one, because there is
no way to turn off the slowly flashing red glow on the video card).&lt;&#x2F;p&gt;
&lt;p&gt;Here is the part-list for the build:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;CPU: Intel Core i9-9900k 5.0 GHz&lt;&#x2F;li&gt;
&lt;li&gt;CPU Cooler: Noctua NH-D15 82.5 CFM CPU Cooler&lt;&#x2F;li&gt;
&lt;li&gt;Motherboard: Asus Prime Z390-A ATX LGA1151&lt;&#x2F;li&gt;
&lt;li&gt;Memory: Corsair Vengeance LPX 16 GB (2 x 8 GB) DDR4-3000&lt;&#x2F;li&gt;
&lt;li&gt;Storage: Samsung 970 Evo 250 GB M.2-2280 Solid State Drive&lt;&#x2F;li&gt;
&lt;li&gt;Video Card: Asus GeForce GTX 1070 8 GB STRIX Video Card&lt;&#x2F;li&gt;
&lt;li&gt;Case: Corsair 760T Black ATX Full Tower Case&lt;&#x2F;li&gt;
&lt;li&gt;Power Supply: EVGA SuperNOVA G2 650 W 80+ Gold Cerified Fully-Modular ATX&lt;&#x2F;li&gt;
&lt;li&gt;Wireless Adapter: Gigbyte GC-WB867D-I PCI-Express x1 802.11a&#x2F;b&#x2F;g&#x2F;n&#x2F;ac WiFi
Adapter&lt;&#x2F;li&gt;
&lt;li&gt;Thunderbolt AIC: Asus ThunderboltEX 3&lt;&#x2F;li&gt;
&lt;li&gt;Operating System: Ubuntu Linux 18.10&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;After building the system with this configuration, running the Nvidia&#x27;s DP
output into the Thunderbolt AIC, and connecting the LG 5k monitor via its
Thunderbolt cable, a fresh install of Ubuntu loaded up fine, and we were able to
log into the Gnome 3 desktop just fine. Unfortunately (or perhaps fortunately in
this case?), Nvidia does not provide their own open-source drivers for Linux,
and Ubuntu does not install the closed-source drivers by default. Instead, you
get the open-source nouveau driver. That driver works fine for most things, but
if we wanted to settle for &quot;most things&quot;, we&#x27;d just have used the on-board GPU
of the motherboard. In order to take advantage of the Nvidia&#x27;s capabilities, we
needed to install the proprietary Nvidia drivers. Ubuntu&#x27;s apt repository
contains the &lt;code&gt;nvidia-driver-390&lt;&#x2F;code&gt; package, and they even make it easy to switch
to it by running &lt;code&gt;ubuntu-drivers autoinstall&lt;&#x2F;code&gt;, however the distribution lags
behind the official release of the drivers, which are up to version 415 as of
this writing. In order to have access to the latest drivers, just add the PPA
and update apt:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;sudo add-apt-repository ppa:graphics-drivers&#x2F;ppa
sudo apt-get update
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Then you can install the proprietary driver with:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;sudo apt install nvidia-driver-415 # or whatever the latest version is
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This will add the appropriate kernel module and some necessary configuration.&lt;&#x2F;p&gt;
&lt;p&gt;Then you&#x27;ll reboot your computer, and...&lt;&#x2F;p&gt;
&lt;p&gt;The desktop manager won&#x27;t load. You&#x27;ll probably be staring at a blank screen,
perhaps with a blinking cursor.&lt;&#x2F;p&gt;
&lt;p&gt;Luckily, I do have another display here at the house that has an HDMI input, so
we were able to hook that display up for some debugging. We hooked up the other
display and rebooted. This time, GDM loaded up on the HDMI display just fine. I
then reattached the LG display as a second monitor. Nothing showed up on it,
&lt;em&gt;but&lt;&#x2F;em&gt;, it showed up as a monitor in Gnome&#x27;s display settings. Heh. That&#x27;s
weird...&lt;&#x2F;p&gt;
&lt;p&gt;On a whim, I changed the resolution of the LG display from &quot;auto&quot; to 4096x2304,
and...&lt;em&gt;all of the sudden it worked!&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;I don&#x27;t know for sure, but my guess is that something about being passed through
the Thunderbolt AIC is causing the automatic detection of the resolution to
fail.&lt;&#x2F;p&gt;
&lt;p&gt;Great, we&#x27;re done now, right?&lt;&#x2F;p&gt;
&lt;p&gt;Not exactly. I&#x27;ll spare you the details of the rest of the debugging we went
through (&lt;em&gt;lot&#x27;s&lt;&#x2F;em&gt; of trial and error, searching Google, and piecing together what
I could find from people who had sort-of-kind-of similar symptoms here and
there, but never all in the same combination, and definitely never with this set
of hardware.) Suffice it to say that the solution lies in being &lt;em&gt;very&lt;&#x2F;em&gt; specific
about configuring the display with all three of Xorg, GDM, and Gnome. Here are
the contents of the configuration files that needed to be changed in order to
get this working on the new system (you may need to tweak them a bit if you have
a different combination of hardware, but hopefully this saves you from the hours
of trial and error!):&lt;&#x2F;p&gt;
&lt;p&gt;In &lt;code&gt;&#x2F;etc&#x2F;X11&#x2F;xorg.conf&lt;&#x2F;code&gt;:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;# nvidia-settings: X configuration file generated by nvidia-settings
# nvidia-settings:  version 415.27


Section &amp;quot;ServerLayout&amp;quot;
    Identifier     &amp;quot;Layout0&amp;quot;
    Screen      0  &amp;quot;Screen0&amp;quot; 0 0
    InputDevice    &amp;quot;Keyboard0&amp;quot; &amp;quot;CoreKeyboard&amp;quot;
    InputDevice    &amp;quot;Mouse0&amp;quot; &amp;quot;CorePointer&amp;quot;
    Option         &amp;quot;Xinerama&amp;quot; &amp;quot;0&amp;quot;
EndSection

Section &amp;quot;Files&amp;quot;
EndSection

Section &amp;quot;Module&amp;quot;
    Load           &amp;quot;dbe&amp;quot;
    Load           &amp;quot;extmod&amp;quot;
    Load           &amp;quot;type1&amp;quot;
    Load           &amp;quot;freetype&amp;quot;
    Load           &amp;quot;glx&amp;quot;
EndSection

Section &amp;quot;InputDevice&amp;quot;

    # generated from default
    Identifier     &amp;quot;Mouse0&amp;quot;
    Driver         &amp;quot;mouse&amp;quot;
    Option         &amp;quot;Protocol&amp;quot; &amp;quot;auto&amp;quot;
    Option         &amp;quot;Device&amp;quot; &amp;quot;&#x2F;dev&#x2F;psaux&amp;quot;
    Option         &amp;quot;Emulate3Buttons&amp;quot; &amp;quot;no&amp;quot;
    Option         &amp;quot;ZAxisMapping&amp;quot; &amp;quot;4 5&amp;quot;
EndSection

Section &amp;quot;InputDevice&amp;quot;

    # generated from default
    Identifier     &amp;quot;Keyboard0&amp;quot;
    Driver         &amp;quot;kbd&amp;quot;
EndSection

Section &amp;quot;Monitor&amp;quot;

    # HorizSync source: edid, VertRefresh source: edid
    Identifier     &amp;quot;Monitor0&amp;quot;
    VendorName     &amp;quot;Unknown&amp;quot;
    ModelName      &amp;quot;LG Electronics LG UltraFine&amp;quot;
    HorizSync       31.5 - 177.7
    VertRefresh     60.0 - 60.0
    Option         &amp;quot;DPMS&amp;quot;
EndSection

Section &amp;quot;Device&amp;quot;
    Identifier     &amp;quot;Device0&amp;quot;
    Driver         &amp;quot;nvidia&amp;quot;
    VendorName     &amp;quot;NVIDIA Corporation&amp;quot;
    BoardName      &amp;quot;GeForce GTX 1070&amp;quot;
EndSection

Section &amp;quot;Screen&amp;quot;

# Removed Option &amp;quot;metamodes&amp;quot; &amp;quot;DP-0: nvidia-auto-select +0+0&amp;quot;
    Identifier     &amp;quot;Screen0&amp;quot;
    Device         &amp;quot;Device0&amp;quot;
    Monitor        &amp;quot;Monitor0&amp;quot;
    DefaultDepth    24
    Option         &amp;quot;Stereo&amp;quot; &amp;quot;0&amp;quot;
    Option         &amp;quot;nvidiaXineramaInfoOrder&amp;quot; &amp;quot;DFP-3&amp;quot;
    Option         &amp;quot;ConnectedMonitor&amp;quot; &amp;quot;DP-0&amp;quot;
    Option         &amp;quot;metamodes&amp;quot; &amp;quot;DP-0: 4096x2304 +0+0&amp;quot;
    Option         &amp;quot;SLI&amp;quot; &amp;quot;Off&amp;quot;
    Option         &amp;quot;MultiGPU&amp;quot; &amp;quot;Off&amp;quot;
    Option         &amp;quot;BaseMosaic&amp;quot; &amp;quot;off&amp;quot;
    SubSection     &amp;quot;Display&amp;quot;
        Depth       24
    EndSubSection
EndSection
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;In &lt;em&gt;both&lt;&#x2F;em&gt; &lt;code&gt;$HOME&#x2F;.config&#x2F;monitors.xml&lt;&#x2F;code&gt; &lt;em&gt;and&lt;&#x2F;em&gt;
&lt;code&gt;&#x2F;var&#x2F;lib&#x2F;gdm3&#x2F;.config&#x2F;monitors.xml&lt;&#x2F;code&gt; (my guess is that you&#x27;ll need to change the
value for &lt;code&gt;&amp;lt;serial&amp;gt;&lt;&#x2F;code&gt; here to the serial number for your actual monitor):&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;xml&quot;&gt;&amp;lt;monitors version=&amp;quot;2&amp;quot;&amp;gt;
  &amp;lt;configuration&amp;gt;
    &amp;lt;logicalmonitor&amp;gt;
      &amp;lt;x&amp;gt;0&amp;lt;&#x2F;x&amp;gt;
      &amp;lt;y&amp;gt;0&amp;lt;&#x2F;y&amp;gt;
      &amp;lt;scale&amp;gt;1&amp;lt;&#x2F;scale&amp;gt;
      &amp;lt;primary&amp;gt;yes&amp;lt;&#x2F;primary&amp;gt;
      &amp;lt;monitor&amp;gt;
        &amp;lt;monitorspec&amp;gt;
          &amp;lt;connector&amp;gt;DP-0&amp;lt;&#x2F;connector&amp;gt;
          &amp;lt;vendor&amp;gt;GSM&amp;lt;&#x2F;vendor&amp;gt;
          &amp;lt;product&amp;gt;LG UltraFine&amp;lt;&#x2F;product&amp;gt;
          &amp;lt;serial&amp;gt;707NTFA2N827&amp;lt;&#x2F;serial&amp;gt;
        &amp;lt;&#x2F;monitorspec&amp;gt;
        &amp;lt;mode&amp;gt;
          &amp;lt;width&amp;gt;4096&amp;lt;&#x2F;width&amp;gt;
          &amp;lt;height&amp;gt;2304&amp;lt;&#x2F;height&amp;gt;
          &amp;lt;rate&amp;gt;59.999275207519531&amp;lt;&#x2F;rate&amp;gt;
        &amp;lt;&#x2F;mode&amp;gt;
      &amp;lt;&#x2F;monitor&amp;gt;
    &amp;lt;&#x2F;logicalmonitor&amp;gt;
  &amp;lt;&#x2F;configuration&amp;gt;
&amp;lt;&#x2F;monitors&amp;gt;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;With those files in place, neither X nor Gnome were relying on autoconfiguration
of the display&#x27;s settings, and the LG 5k display seems to be working just fine
(although not actually at &quot;5k&quot;, but it still looks great.)&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Complex Unique Constraints with PostgreSQL Triggers in Ecto</title>
          <pubDate>Sun, 16 Feb 2020 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/complex-unique-constraints-with-postgresql-triggers-in-ecto/</link>
          <guid>https://johnwilger.com/blog/complex-unique-constraints-with-postgresql-triggers-in-ecto/</guid>
          <description xml:base="https://johnwilger.com/blog/complex-unique-constraints-with-postgresql-triggers-in-ecto/">&lt;p&gt;&lt;a rel=&quot;external&quot; title=&quot;Ecto&quot; href=&quot;https:&#x2F;&#x2F;hexdocs.pm&#x2F;ecto&quot;&gt;Ecto&lt;&#x2F;a&gt; makes it easy to work with typical uniqueness constraints in your database; you just
define your table like this:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule MyApp.Repo.Migrations.CreateFoos do
  use Ecto.Migration

  def change do
    create table(:foos) do
      add :name, :text, null: false
    end

    create unique_index(:foos, :name)
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;and a module with a &lt;a rel=&quot;external&quot; title=&quot;Ecto.Changeset.unique_constraint&#x2F;3&quot; href=&quot;https:&#x2F;&#x2F;hexdocs.pm&#x2F;ecto&#x2F;Ecto.Changeset.html#unique_constraint&#x2F;3&quot;&gt;changeset validation for the uniqueness constraint&lt;&#x2F;a&gt;,
perhaps like:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule MyApp.Foo do
  use Ecto.Schema

  import Ecto.Changeset

  schema &amp;quot;foos&amp;quot; do
    field :name, :string
  end

  def changeset(%__MODULE__{} = foo, %{} = changes) do
      foo
      |&amp;gt; cast(changes, [:name])
      |&amp;gt; unique_constraint(:name)
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Then, when you run:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;result = MyApp.Foo.changeset(%MyApp.Foo{}, %{name: &amp;quot;bar&amp;quot;})
         |&amp;gt; MyApp.Repo.insert()
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The Ecto library will attempt to insert your record, and if there is already a record where the
name column is set to &quot;bar&quot;, Ecto will see the uniqueness constraint violation error produced by
the database and turn it into a validation error on your changeset, so that it would look
something like:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;{:error, %Ecto.Changeset{errors: [name: {&amp;quot;has already been taken&amp;quot;, _}]}} = result
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;something-more-complex&quot;&gt;Something More Complex&lt;&#x2F;h3&gt;
&lt;p&gt;I recently needed to enforce a database constraint similar in spirit to a unique index (if a
record were in violation, I wanted the same behavior on the Elixir end of things—the changeset
should report that the value for the field &quot;has already been taken&quot;) however the criteria for what
should be considered &quot;unique&quot; was more complex than what a simple unique index in PostgreSQL would
be able to deal with.&lt;&#x2F;p&gt;
&lt;p&gt;Let&#x27;s pretend we have a project that is managing meeting room reservations. The process for
reserving a meeting room is as follows:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;The customer chooses a room and selects the time range for the reservation.&lt;&#x2F;li&gt;
&lt;li&gt;Assuming the room is available at that time, the system places a hold on the room for 24 hours&lt;&#x2F;li&gt;
&lt;li&gt;Within that 24-hour period, the customer pays for the room and completes some contractual
information that must be signed by both the customer and the facility manager.&lt;&#x2F;li&gt;
&lt;li&gt;Once the payment is complete and the contracts are signed, the reservation is confirmed.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;Here, then, are the business rules for creating a reservation:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A room reservation cannot be made for a given room and time-period if there exists another room
reservation for that same room with an overlapping time-period and the existing reservation is
in a &quot;confirmed&quot; status.&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;A room reservation cannot be made for a given room and time-period if there exists another room
reservation for that same room with an overlapping time-period, the existing reservation is in a
&quot;hold&quot; status, and the existing reservation was created within the last 24 hours.&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;We decide to store the room reservations in a table defined as:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule Meetings.Repo.Migrations.CreateRoomReservations do
  use Ecto.Migration

  def change do
    create table(:room_reservations) do
      add :customer_id, :binary_id, null: false
      add :room_id, :binary_id, null: false
      add :reservation_starts_at, :timestamp, null: false
      add :reservation_ends_at, :timestamp, null: false
      add :status, :text, null: false, default: &amp;quot;hold&amp;quot;
      timestamps()
    end
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;and ideally, we&#x27;d like to be able to model the &lt;code&gt;RoomReservation&lt;&#x2F;code&gt; in Elixir as:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule Meetings.RoomReservation do
  use Ecto.Schema

  import Ecto.Changeset

  schema &amp;quot;room_reservations&amp;quot; do
    field :customer_id, :binary_id
    field :room_id, :binary_id
    field :reservation_starts_at, :utc_datetime
    field :reservation_ends_at, :utc_datetime
    field :status, :string
    timestamps()
  end

  def create(%{} = res_data) do
    %__MODULE__{}
    |&amp;gt; cast(res_data, [:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
    |&amp;gt; validate_required([:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
    |&amp;gt; put_change(:status, &amp;quot;hold&amp;quot;)
    |&amp;gt; unique_constraint(:room_id, message: &amp;quot;has already been reserved&amp;quot;)
    |&amp;gt; Meeting.Repo.insert()
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;with the interesting bit being the &lt;code&gt;|&amp;gt; unique_constraint(:room_id)&lt;&#x2F;code&gt; line. When a customer tries to
reserve a room that is already reserved for the given time period, we want the
&lt;code&gt;Meeting.RoomReservation.create&#x2F;1&lt;&#x2F;code&gt; function to return with a validation error:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;%{:error, %Ecto.Changeset{errors: [room_id: {&amp;quot;has already been reserved&amp;quot;, _}]}} =
  Meetings.RoomReservation.create(%{
    customer_id: &amp;quot;0af5716a-c3d2-4d4c-87f5-9fed9b2515d4&amp;quot;,
    room_id: &amp;quot;2c0d6242-3585-40cb-be92-a1818c0f4a73&amp;quot;,
    reservation_starts_at: DateTime.from_naive!(~N[2020-01-01 8:00:00], &amp;quot;Etc&#x2F;UTC&amp;quot;),
    reservation_ends_at: DateTime.from_naive!(~N[2020-01-01 17:00:00], &amp;quot;Etc&#x2F;UTC&amp;quot;)
  })
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Clearly, according to the business rules for the system, we can&#x27;t just put a unique index on the
&lt;code&gt;room_id&lt;&#x2F;code&gt; column. To the best of my knowledge, an exclusion constraint also won&#x27;t work here,
because of the need to &lt;em&gt;not&lt;&#x2F;em&gt; check against any rows that have a status of &quot;hold&quot; and an
&lt;code&gt;inserted_at&lt;&#x2F;code&gt; timestamp that is more than 24 hours ago relative to the current time. It &lt;em&gt;is&lt;&#x2F;em&gt;
however possible to use a trigger function to enforce the constraint in the database:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule Meetings.Repo.Migrations.CreateRoomReservations do
  use Ecto.Migration

  def change do
    create table(:room_reservations) do
      add :customer_id, :binary_id, null: false
      add :room_id, :binary_id, null: false
      add :reservation_starts_at, :timestamp, null: false
      add :reservation_ends_at, :timestamp, null: false
      add :status, :text, null: false, default: &amp;quot;hold&amp;quot;
      timestamps()
    end

    execute(
      # up
      ~S&amp;quot;&amp;quot;&amp;quot;
      CREATE FUNCTION room_reservations_check_room_availability() RETURNS TRIGGER AS
      $$
      BEGIN
      IF EXISTS(
        SELECT 1
        FROM room_reservations rr
        WHERE rr.room_id = NEW.room_id
          AND tsrange(rr.reservation_starts_at, rr.reservation_ends_at, &amp;#39;[]&amp;#39;) &amp;amp;&amp;amp;
              tsrange(NEW.reservation_starts_at, NEW.reservation_ends_at, &amp;#39;[]&amp;#39;)
          AND (
            rr.status = &amp;#39;confirmed&amp;#39;
              OR rr.inserted_at &amp;gt; CURRENT_TIMESTAMP - interval &amp;#39;24 hours&amp;#39;
          )
      )
      THEN
        RAISE &amp;quot;room already reserved&amp;quot;;
      END IF;
      RETURN NEW;
      END;
      $$ language plpgsql;
      &amp;quot;&amp;quot;&amp;quot;,

      # down
      &amp;quot;DROP FUNCTION room_reservations_check_room_availability;&amp;quot;
    )

    execute(
      # up
      ~S&amp;quot;&amp;quot;&amp;quot;
      CREATE TRIGGER room_reservations_room_availability_check
      BEFORE INSERT ON room_reservations
      FOR EACH ROW
      EXECUTE PROCEDURE room_reservations_check_room_availability();
      &amp;quot;&amp;quot;&amp;quot;,

      # down
      &amp;quot;DROP TRIGGER room_reservations_room_availability_check ON room_reservations;&amp;quot;
    )
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The problem is that instead of that nice validation error on the changeset that we &lt;em&gt;want&lt;&#x2F;em&gt; to get,
we instead end up with an unhandled &lt;code&gt;Postgrex.Error&lt;&#x2F;code&gt; exception. We could, of course, simply rescue
that exception in our &lt;code&gt;Meetings.RoomReservation.create&#x2F;1&lt;&#x2F;code&gt; function and add the error to the
changeset ourselves, but that starts to look a little ugly:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule Meetings.RoomReservation do
  # ...

  def create(%{} = res_data) do
    changeset =
      %__MODULE__{}
      |&amp;gt; cast(res_data, [:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
      |&amp;gt; validate_required([:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
      |&amp;gt; put_change(:status, &amp;quot;hold&amp;quot;)
    Meeting.Repo.insert(changeset)
  rescue
    error in Postgrex.Error -&amp;gt;
      case error do
        %{postgres: %{message: &amp;quot;room already reserved&amp;quot;}} -&amp;gt;
          changeset
          |&amp;gt; add_error(:room_id, &amp;quot;has already been reserved&amp;quot;)
      end
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;tl-dr-it-s-all-about-the-raise&quot;&gt;TL;DR - It&#x27;s All About the Raise&lt;&#x2F;h3&gt;
&lt;p&gt;Knowing that &lt;code&gt;Ecto.Changeset.unique_constraint&#x2F;3&lt;&#x2F;code&gt; works by intercepting an error raised by the
database, I set out to see if I could implement the complex unique constraint logic in the
database and still be able to use the &lt;code&gt;Ecto.Changeset.unique_constraint&#x2F;3&lt;&#x2F;code&gt; validation without
needing to modify any Elixir code. Looking at &lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;elixir-ecto&#x2F;ecto_sql&#x2F;blob&#x2F;0359b7ce5155974d566fcd0b127f8695aed8b3a9&#x2F;lib&#x2F;ecto&#x2F;adapters&#x2F;postgres&#x2F;connection.ex#L18-L19&quot;&gt;the relevant code in ecto_sql&lt;&#x2F;a&gt;, we
can see that the trick to getting the &lt;code&gt;room_reservations_check_room_availability&lt;&#x2F;code&gt; functionality to
work with that function is to change the PostgreSQL exception to use the correct error code as
well as a constraint name that is linked with the changeset validation. We can redifine the
PostgreSQL function as:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sql&quot;&gt;CREATE FUNCTION room_reservations_check_room_availability() RETURNS TRIGGER AS
$$
BEGIN
IF EXISTS(
  SELECT 1
  FROM room_reservations rr
  WHERE rr.room_id = NEW.room_id
    AND tsrange(rr.reservation_starts_at, rr.reservation_ends_at, &amp;#39;[]&amp;#39;) &amp;amp;&amp;amp;
        tsrange(NEW.reservation_starts_at, NEW.reservation_ends_at, &amp;#39;[]&amp;#39;)
    AND (
      rr.status = &amp;#39;confirmed&amp;#39;
        OR rr.inserted_at &amp;gt; CURRENT_TIMESTAMP - interval &amp;#39;24 hours&amp;#39;
    )
)
THEN
  RAISE unique_violation
    USING CONSTRAINT = &amp;#39;room_reservations_room_reserved&amp;#39;;
END IF;
RETURN NEW;
END;
$$ language plpgsql;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;and then add the &lt;code&gt;:name&lt;&#x2F;code&gt; option to our &lt;code&gt;unique_constraint&lt;&#x2F;code&gt; in the changeset validation:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;elixir&quot;&gt;defmodule Meetings.RoomReservation do
  #...

  def create(%{} = res_data) do
    %__MODULE__{}
    |&amp;gt; cast(res_data, [:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
    |&amp;gt; validate_required([:customer_id, :room_id, :reservation_starts_at, :reservation_ends_at])
    |&amp;gt; put_change(:status, &amp;quot;hold&amp;quot;)
    |&amp;gt; unique_constraint(:room_id,
                         message: &amp;quot;has already been reserved&amp;quot;,
                         name: &amp;quot;room_reservations_room_reserved&amp;quot;)
    |&amp;gt; Meeting.Repo.insert()
  end
end
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
</description>
      </item>
      <item>
          <title>Kookaburra 0.24.0 Released - Exorcised ActiveSupport</title>
          <pubDate>Tue, 15 May 2012 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/kookaburra-0240-released-exorcised-activesupport/</link>
          <guid>https://johnwilger.com/blog/kookaburra-0240-released-exorcised-activesupport/</guid>
          <description xml:base="https://johnwilger.com/blog/kookaburra-0240-released-exorcised-activesupport/">&lt;p&gt;I just released version 0.24.0 of &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;kookaburra&quot;&gt;Kookaburra&lt;&#x2F;a&gt; to
&lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;rubygems.org&#x2F;gems&#x2F;kookaburra&#x2F;versions&#x2F;0.24.0&quot;&gt;Rubygems.org&lt;&#x2F;a&gt;. While there were no changes to &lt;em&gt;Kookaburra&#x27;s&lt;&#x2F;em&gt; API
with this release, it is a minor release rather than a patch release,
because I removed &lt;a rel=&quot;external&quot; title=&quot;ActiveSupport 3.0.0 API Documentation&quot; href=&quot;http:&#x2F;&#x2F;rubydoc.info&#x2F;gems&#x2F;activesupport&#x2F;3.0.0&#x2F;frames&quot;&gt;ActiveSupport&lt;&#x2F;a&gt; from Kookaburra&#x27;s dependencies.
Previously, Kookaburra depended on ActiveSupport &amp;gt;= 3.0, and this
prevented it from working smoothly with Rails 2.x applications (at least
in the same Bundler bundle.) Since Kookaburra only used a small fraction
of ActiveSupport, it seemed easiest just to break the dependency, so
that your application can use whatever version of ActiveSupport it
needs.&lt;&#x2F;p&gt;
&lt;p&gt;Because of the fact that ActiveSupport makes a number of modifications
to core Ruby classes, if your application does not otherwise explicitly
require ActiveSupport, some of these core classes may now behave
differently. Specifically, &lt;code&gt;Hash&lt;&#x2F;code&gt; and &lt;code&gt;String&lt;&#x2F;code&gt; core extensions were
being used, as well as &lt;code&gt;ActiveSupport::CoreExt::Module::Delegation&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Also, where &lt;code&gt;Kookaburra::JsonApiDriver&lt;&#x2F;code&gt; used &lt;code&gt;ActiveSupport::JSON&lt;&#x2F;code&gt; for
encoding and decoding of JSON data, it now uses &lt;code&gt;::JSON&lt;&#x2F;code&gt; as provided by
&lt;em&gt;either&lt;&#x2F;em&gt; the &lt;a rel=&quot;external&quot; title=&quot;json | RubyGems.org&quot; href=&quot;http:&#x2F;&#x2F;rubygems.org&#x2F;gems&#x2F;json&quot;&gt;json&lt;&#x2F;a&gt; or the &lt;a rel=&quot;external&quot; title=&quot;json_pure | RubyGems.org&quot; href=&quot;http:&#x2F;&#x2F;rubygems.org&#x2F;gems&#x2F;json_pure&quot;&gt;json_pure&lt;&#x2F;a&gt; gems. The tests
on Kookaburra itself pass just by swapping out ActiveSupport JSON
library for the new one, but these tests do not comprehensively test
JSON usage, and there may be differences that impact your application.&lt;&#x2F;p&gt;
&lt;p&gt;Kookaburra declares its dependency on the &lt;code&gt;json_pure&lt;&#x2F;code&gt; gem (a pure-ruby
version) in order to ensure that it works cross-platform. However, both
&lt;code&gt;json&lt;&#x2F;code&gt; and &lt;code&gt;json_pure&lt;&#x2F;code&gt; are written in such a way that, when you &lt;code&gt;require &#x27;json&#x27;&lt;&#x2F;code&gt;, it will detect if both libraries are installed and prefer the
natively compiled &lt;code&gt;json&lt;&#x2F;code&gt; gem&#x27;s version of the library.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>tmux and the OSX Clipboard</title>
          <pubDate>Thu, 12 Apr 2012 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/tmux-and-the-osx-clipboard/</link>
          <guid>https://johnwilger.com/blog/tmux-and-the-osx-clipboard/</guid>
          <description xml:base="https://johnwilger.com/blog/tmux-and-the-osx-clipboard/">&lt;p&gt;I started using &lt;a rel=&quot;external&quot; title=&quot;tmux&quot; href=&quot;http:&#x2F;&#x2F;tmux.sourceforge.net&#x2F;&quot;&gt;tmux&lt;&#x2F;a&gt; recently after a) I was informed that GNU screen is
basically an outdated POS, and b) an &lt;a rel=&quot;external&quot; title=&quot;tmux: Productive Mouse-Free Development&quot; href=&quot;http:&#x2F;&#x2F;pragprog.com&#x2F;book&#x2F;bhtmux&#x2F;tmux&quot;&gt;excellent book&lt;&#x2F;a&gt; on the subject
was published. I&#x27;m glad to have made the switch, as tmux is a wonderful
improvement over screen. However, on OSX, I found that the &lt;code&gt;pbcopy&lt;&#x2F;code&gt; and
&lt;code&gt;pbpaste&lt;&#x2F;code&gt; commands (among other things) would no longer work from within a tmux
session.&lt;&#x2F;p&gt;
&lt;p&gt;After a bit of searching, I found &lt;a rel=&quot;external&quot; title=&quot;Reintroducing tmux to the OSX Clipboard&quot; href=&quot;http:&#x2F;&#x2F;writeheavy.com&#x2F;2011&#x2F;10&#x2F;23&#x2F;reintroducing-tmux-to-the-osx-clipboard.html&quot;&gt;this post&lt;&#x2F;a&gt; that explains how to make
things right again.&lt;&#x2F;p&gt;
&lt;p&gt;TL;DR:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh run this&quot;&gt;brew install reattach-to-user-namespace
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;pre&gt;&lt;code data-lang=&quot;sh add this to ~&#x2F;.tmux.conf&quot;&gt;# Reattach to user namespace so that clipboard integration with OSX works again
set-option -g default-command &amp;quot;reattach-to-user-namespace -l zsh&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Obviously, replace &lt;code&gt;zsh&lt;&#x2F;code&gt; above with whatever you use for your shell.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Kookaburra Rewrite for 0.15.1</title>
          <pubDate>Sat, 10 Mar 2012 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/kookaburra-rewrite-for-0151/</link>
          <guid>https://johnwilger.com/blog/kookaburra-rewrite-for-0151/</guid>
          <description xml:base="https://johnwilger.com/blog/kookaburra-rewrite-for-0151/">&lt;p&gt;After getting some good feedback on &lt;a rel=&quot;external&quot; title=&quot;jwilger&#x2F;kookaburra&quot; href=&quot;http:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;kookaburra&quot;&gt;Kookaburra&lt;&#x2F;a&gt; &lt;a rel=&quot;external&quot; title=&quot;jwilger&#x2F;kookaburra&quot; href=&quot;http:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;kookaburra&quot;&gt;Kookaburra&lt;&#x2F;a&gt; since the original
release announcement as well as using it in a few more projects, [Sam] &lt;a rel=&quot;external&quot; title=&quot;Sam Livingston-Gray&quot; href=&quot;http:&#x2F;&#x2F;livingston-gray.com&quot;&gt;SLG&lt;&#x2F;a&gt; and
I decided to treat all of the versions prior to 0.15.0 as a spike and rewrite
the framework from the ground up. Although we certainly learned a lot about the
approach with the pre-0.15 versions, the problems with continuing to grow
Kookaburra from that seed became apparent as we tried to use it in more
applications.&lt;&#x2F;p&gt;
&lt;p&gt;Because the original code was extracted from another project, the Kookaburra
framework itself was not well tested. It had been indirectly tested due to its
position within that other project&#x27;s tests, but as a seperate library it lacked
tests that would help bot document the project and ensure newcomers can work on
it with a good safety net. Additionally, there were a number of design decisions
made in the earlier versions of Kookaburra that were...questionable. Not only
would these decisions probably have been made differently if the development of
the framework itself had been driven with unit tests, these decisions in many
cases made it difficult to go back and add tests after the fact.&lt;&#x2F;p&gt;
&lt;p&gt;Starting with 0.15.1, which I just [released] &lt;a rel=&quot;external&quot; title=&quot;kookaburra | Rubygems.org&quot; href=&quot;http:&#x2F;&#x2F;rubygems.org&#x2F;gems&#x2F;kookaburra&quot;&gt;Gem&lt;&#x2F;a&gt; a few moments ago,
Kookaburra has been rewritten using a proper TDD approach. In some ways, 0.15.1
is possibly less functional than its 0.14.x cousins, but it is almost certainly
easier to understand and contribute to the project from this point on. Please
have a look at the new codebase, try it out in your projects, and send me a pull
request with updates to the library. Also, if you have any problems getting it
running, please don&#x27;t hesitate to [let me know] &lt;a rel=&quot;external&quot; title=&quot;Issues jwilger&#x2F;kookaburra&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jwilger&#x2F;kookaburra&#x2F;issues?sort=created&amp;amp;direction=desc&amp;amp;state=open&quot;&gt;GHIssues&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
</description>
      </item>
      <item>
          <title>Using Jeweler for Private Gems</title>
          <pubDate>Wed, 25 Jan 2012 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/using-jeweler-for-private-gems/</link>
          <guid>https://johnwilger.com/blog/using-jeweler-for-private-gems/</guid>
          <description xml:base="https://johnwilger.com/blog/using-jeweler-for-private-gems/">&lt;p&gt;I really like using [Jeweler] &lt;a rel=&quot;external&quot; title=&quot;technicalpickles&#x2F;jeweler - GitHub&quot; href=&quot;http:&#x2F;&#x2F;github.com&#x2F;technicalpickles&#x2F;jeweler&quot;&gt;1&lt;&#x2F;a&gt; to create and manage gems, but its default
behavior is to publish your gem to [rubygems.org] &lt;a rel=&quot;external&quot; title=&quot;RubyGems.org | your community gem host&quot; href=&quot;http:&#x2F;&#x2F;rubygems.org&quot;&gt;2&lt;&#x2F;a&gt; whenever you run &lt;code&gt;rake release&lt;&#x2F;code&gt;. This is great for generally useful libraries that you want to
open-source, but not as great when you want to use gems as shared libraries for
internal use only (whether because the code contains business secrets or just
because it&#x27;s not something that would be useful to the community at large.)&lt;&#x2F;p&gt;
&lt;p&gt;After a bit of poking around, I figured out how to use Jeweler to manage
your versioning, gemspec and git tagging without pushing the result to
rubygems.org when running the release task. It turns out to be both obvious and
trivial. Just add the following line to the end of your gem project&#x27;s
&lt;code&gt;Rakefile&lt;&#x2F;code&gt;:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code data-lang=&quot;ruby Rakefile&quot;&gt;Rake::Task[:release].prerequisites.delete(&amp;#39;gemcutter:release&amp;#39;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
</description>
      </item>
      <item>
          <title>Acceptance and Integration Testing with Kookaburra</title>
          <pubDate>Sat, 21 Jan 2012 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/acceptance-and-integration-testing-with-kookaburra/</link>
          <guid>https://johnwilger.com/blog/acceptance-and-integration-testing-with-kookaburra/</guid>
          <description xml:base="https://johnwilger.com/blog/acceptance-and-integration-testing-with-kookaburra/">&lt;p&gt;&lt;strong&gt;UPDATE (2012-01-22):&lt;&#x2F;strong&gt; &lt;em&gt;I realized this morning that the credit I gave to &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;resume.livingston-gray.com&#x2F;&quot;&gt;Sam
Livingston-Gray&lt;&#x2F;a&gt; below may not have
adequately shown how instrumental he was in getting this project off the ground;
especially since much of his work was done in the private repository from which
this was extracted. So, thanks, Sam. This might not have gone anywhere if you
hadn&#x27;t worked to put the idea in practice in our application and helped everyone
on our team learn how to use the approach. I made a few minor changes below to
reflect this a bit better.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;We&#x27;ve been using &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;cukes.info&#x2F;&quot;&gt;Cucumber&lt;&#x2F;a&gt; for acceptance testing at
&lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;renewfund.com&quot;&gt;Renewable Funding&lt;&#x2F;a&gt; since back when it was still part of
the &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;rspec.info&quot;&gt;RSpec&lt;&#x2F;a&gt; project (indeed, since before we were even
Renewable Funding). While we&#x27;ve always liked the ability to have plain-language
feature documentation that we could automatically test against, after years of
adding to and maintaining a fairly large set of Cucumber scenarios, the cost of
that maintenance was starting to really slow us down. The test suite began to
grow fragile, and it seemed like every time one of our UX designers changed
anything about the application&#x27;s interface, the development team would spend a
bunch of time just babysitting Cucumber tests to get them passing again.&lt;&#x2F;p&gt;
&lt;p&gt;Last year, as I was reading Jez Humble&#x27;s excellent &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;www.amazon.com&#x2F;gp&#x2F;product&#x2F;0321601912?tag=contindelive-20&quot;&gt;Continuous
Delivery&lt;&#x2F;a&gt; book,
I was inspired when I came across the section titled &quot;The Application Driver
Layer&quot; (p. 198). This section describes an approach to acceptance testing where
the specification and the test implementation are isolated from the details of
the application&#x27;s user interface by inserting a layer between the two that uses
good old OOP to abstract the user interface components. Martin Fowler describes
it as the [Window Driver] &lt;a rel=&quot;external&quot; title=&quot;Window Driver - Martin Fowler&quot; href=&quot;http:&#x2F;&#x2F;martinfowler.com&#x2F;eaaDev&#x2F;WindowDriver.html&quot;&gt;1&lt;&#x2F;a&gt; pattern on his website.&lt;&#x2F;p&gt;
&lt;p&gt;I started a proof-of-concept implementation of this pattern last summer, then my
coworker, &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;resume.livingston-gray.com&#x2F;&quot;&gt;Sam Livingston-Gray&lt;&#x2F;a&gt; and I
started pulling it into a new project at work. After Sam and the rest of the
Renewable Funding team helped improve on my original attempt while putting it to
use for the last six or so months, we extracted a library to make it easier for
other Ruby developers to implement this pattern in their testing. I&#x27;m happy to
introduce &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;github.com&#x2F;projectdx&#x2F;kookaburra&quot;&gt;Kookaburra&lt;&#x2F;a&gt; to the world.&lt;&#x2F;p&gt;</description>
      </item>
      <item>
          <title>Capybara, Selenium and Firefox 3.5</title>
          <pubDate>Sat, 30 Apr 2011 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/capybara-selenium-and-firefox-35/</link>
          <guid>https://johnwilger.com/blog/capybara-selenium-and-firefox-35/</guid>
          <description xml:base="https://johnwilger.com/blog/capybara-selenium-and-firefox-35/">&lt;p&gt;I banged my head against this one while setting up our CI build for a new
project at work and just now figured it out. I added an example integration
spec that just visits the default Rails homepage, clicks the &quot;About your
application&#x27;s environment&quot; link, and verifies the Rails version. This link
uses an AJAX action to load the environment details, so I marked it as a
JavaScript test, which would cause it to use Capybara&#x27;s Selenium driver.&lt;&#x2F;p&gt;
&lt;p&gt;The test passed fine on my machine (with Firefox 4 installed), but it failed
consistently on CI, which has Firefox 3.5. It kept failing with the exception
&lt;code&gt;Capybara::TimeoutError: failed to resynchronize, AJAX request timed out&lt;&#x2F;code&gt;. I
didn&#x27;t dig enough into the guts of Selenium&#x2F;Firefox to find out exactly what
is going on (and a Google search for that error message turned up nothing that
really explained it) but I did at least find a workable fix. Just add the
following line to &lt;code&gt;spec&#x2F;support&#x2F;capybara.rb&lt;&#x2F;code&gt;:&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;Capybara::Selenium::Driver::DEFAULT_OPTIONS[:resynchronize] = false
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;</description>
      </item>
      <item>
          <title>Apprenticeship Program at Renewable Funding - Thoughts?</title>
          <pubDate>Fri, 04 Mar 2011 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/apprenticeship-program-at-renewable-funding-thoughts/</link>
          <guid>https://johnwilger.com/blog/apprenticeship-program-at-renewable-funding-thoughts/</guid>
          <description xml:base="https://johnwilger.com/blog/apprenticeship-program-at-renewable-funding-thoughts/">&lt;p&gt;The &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;renewfund.com&quot;&gt;Renewable Funding&lt;&#x2F;a&gt; technology team is moving into
some new office space in a few months, and now that we&#x27;ll have our own space,
I&#x27;d like to start up a variation on an apprenticeship program at work.  I&#x27;m
still in the early stages of designing this, and I want to gather some input
from the community before deciding how we might run the program.&lt;&#x2F;p&gt;</description>
      </item>
      <item>
          <title>Production Release Workflow with Git</title>
          <pubDate>Sat, 08 Jan 2011 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/production-release-workflow-with-git/</link>
          <guid>https://johnwilger.com/blog/production-release-workflow-with-git/</guid>
          <description xml:base="https://johnwilger.com/blog/production-release-workflow-with-git/">&lt;p&gt;After growing the &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;www.projectdx.com&quot;&gt;ProjectDX&lt;&#x2F;a&gt; team from three to
eight software developers, our release process was a complete pain, and it
typically took two to three hours to get a good build on the production branch
(and even then some insidious issues would sneak through). By making a few
changes to our development and acceptance process, we were able to turn it
into a five-minute, low-stress job.&lt;&#x2F;p&gt;</description>
      </item>
      <item>
          <title>What it Really Means to be &quot;Agile</title>
          <pubDate>Wed, 15 Dec 2010 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/what-it-really-means-to-be-agile/</link>
          <guid>https://johnwilger.com/blog/what-it-really-means-to-be-agile/</guid>
          <description xml:base="https://johnwilger.com/blog/what-it-really-means-to-be-agile/">&lt;p&gt;Yesterday, Elizabeth Hendrickson posted her &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;testobsessed.com&#x2F;2010&#x2F;12&#x2F;14&#x2F;the-agile-acid-test&#x2F;&quot;&gt;Agile Acid
Test&lt;&#x2F;a&gt; in which she asks
three questions to determine whether or not a team is truly &quot;agile&quot;. There is
also the &lt;a rel=&quot;external&quot; href=&quot;http:&#x2F;&#x2F;agilemanifesto.org&#x2F;&quot;&gt;Agile Manifesto&lt;&#x2F;a&gt; which describes the
values that an agile team should adhere to. While there is nothing that I
disagree with in either Elizabeth&#x27;s post or the Agile Manifesto, there
are two simpler questions you can ask that get to the heart of whether or
not your team is agile: &lt;em&gt;Can you react immediately and without panic when
external constraints on your project are changed&lt;&#x2F;em&gt;, and &lt;em&gt;does your team regularly
and frequently review its processes to ensure the answer to the previous
question is always &quot;yes&quot;?&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;</description>
      </item>
      <item>
          <title>Retrospective Facilitation</title>
          <pubDate>Tue, 13 Jan 2009 00:00:00 +0000</pubDate>
          <author>Unknown</author>
          <link>https://johnwilger.com/blog/retrospective-facilitation/</link>
          <guid>https://johnwilger.com/blog/retrospective-facilitation/</guid>
          <description xml:base="https://johnwilger.com/blog/retrospective-facilitation/">&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Sprint 1 (11&#x2F;03)&lt;&#x2F;th&gt;&lt;th&gt;Sprint 2 (11&#x2F;24)&lt;&#x2F;th&gt;&lt;th&gt;Alpha (12&#x2F;15)&lt;&#x2F;th&gt;&lt;th&gt;Beta (12&#x2F;29)&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;11&#x2F;03&lt;&#x2F;td&gt;&lt;td&gt;11&#x2F;24&lt;&#x2F;td&gt;&lt;td&gt;12&#x2F;15&lt;&#x2F;td&gt;&lt;td&gt;12&#x2F;29&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;I facilitated my second release retrospective with the ProjectDX team yesterday. The first one, back in October, went reasonably well considering I was asked to facilitate at the start of the retrospective and had no time to prepare. However, yesterday&#x27;s retrospective went even better since I knew in advance that I would be facilitating.&lt;&#x2F;p&gt;
&lt;p&gt;The activities we used during the retrospective primarily came from Diana Larsen and Esther Derby&#x27;s book, Agile Retrospectives. We began with the &quot;ESVP&quot; (Explorer, Shopper, Vacationer, Prisoner) activity to gauge the group&#x27;s interest level concerning the work we were doing in the retrospective. Almost everyone in the group identified as an Explorer through an anonymous vote. With this group, I generally trust the outcome, although I did wonder if anyone might have provided the answer they thought they were &quot;supposed&quot; to give.&lt;&#x2F;p&gt;
&lt;p&gt;Next, I asked the team to construct a timeline of events from the past release. Before the retrospective began, I drew the timeline on our whiteboard and divided it into sections for each iteration (our cycle consists of two 3-week sprints followed by a 2-week internal Alpha testing period and a 2-week customer Beta testing period). The result looked something like:&lt;&#x2F;p&gt;
&lt;p&gt;During the retrospective, I asked everyone to take some time to think about any events during that time period that were meaningful to them in any way, write the event on a sticky note, and then place the note on the timeline in approximately the right spot. We spent about 45 minutes generating events, during which time we were free to look back at calendars and email, as well as the commit logs from our SCM. I set a 45-minute time box and mentioned that it was fine for people to take a break if they felt they couldn&#x27;t remember any more events, asking everyone to simply be back on time for the next step.&lt;&#x2F;p&gt;
&lt;p&gt;Instead of specifying that people work in pairs, I chose to let the team self-organize. It was interesting to see how, at first, everyone seemed to work on their own, but as the more obvious events were posted to the board, people started working in groups to come up with more data. At the end of the allotted time, we had a whiteboard full of events.&lt;&#x2F;p&gt;
&lt;p&gt;The area below the timeline was intended for a graph of energy&#x2F;emotion levels, which I&#x27;ll discuss in a moment, but JD Huntington had a great idea to also add a rough bar graph of our commit volume, as reported by GitHub&#x27;s activity graphs, to that area.&lt;&#x2F;p&gt;
&lt;p&gt;After everyone had a last opportunity to look at the board and add any last-minute stickies, I asked everyone to take a turn at the board and use markers to place a blue dot next to any event which they felt positive about and a red dot next to any event that they felt negative about. Then, I asked them to think about their positive and negative energy&#x2F;emotion levels throughout the release and draw a line graph in the bottom section to signify their ups and downs throughout the release cycle. At this point, we looked at each iteration in the cycle and I asked the team to make observations about the data. Using bullet lists at the bottom of the white board, we captured the observations as further data points for analysis.&lt;&#x2F;p&gt;
&lt;p&gt;After a break for lunch, we used the &quot;Patterns &amp;amp; Shifts&quot; activity to generate insights about the data on the time line. We gathered around the white board where I asked each person to mark on the time line any point at which they felt there was a shift or transition of some sort and to draw lines between any events, graph points and observations that they felt were somehow related.&lt;&#x2F;p&gt;
&lt;p&gt;After no one was able to see any more connections, I asked the team to look for any patterns among the connections and the shifts. The first thing we noticed was that we felt like the real shift between iterations was constantly happening about 3 days after the actual time box ended. Rather than get into a long discussion about it when it was noticed, I asked the team to simply keep it in mind as a possible concern during the next phase of the retrospective.&lt;&#x2F;p&gt;
&lt;p&gt;The next observation was that we had a number of connections with long lines spanning two or more iterations. I asked everyone to focus on these connections and look for those where we may have noticed an issue early on but it wasn&#x27;t addressed until much later. In some cases it turned out that while that was the issue, the length of the line just signified that either the work took that long or there was a conscious decision not to address the issue right away.&lt;&#x2F;p&gt;
&lt;p&gt;Reflecting on the other cases led to an interesting discovery, however. Chris noticed that where the lines crossed the bounds of the iterations it was likely signifying a parallel cycle which we were attempting to force into our release cycle: specifically, we have both an explicit product release cycle and an inexplicit customer deployment cycle occurring at the same time.&lt;&#x2F;p&gt;
&lt;p&gt;After we were finished recognizing patterns in the data, it was time to discuss these patterns (and anything else on our minds) and come up with a few action items for the next release cycle. Here I ran the team through the &quot;Circle of Questions&quot; activity. In this activity, the group sits in a circle, and each person takes turns asking a question to the person on their immediate left. The question can be about anything they like (barring anything offensive or attacking), but it&#x27;s helpful to focus on the insights gained during the previous stage of the retrospective. The person to the left answers the question to the best of their ability, and then they ask the person to their left any other question (or the same question if they feel they&#x27;d like a better answer). This continues until the alloted time is up or you have gone around the entire circle twice, whichever comes last. Make sure you go around the complete circle: if some people in the group get more turns to ask or answer a question than others, it can send the wrong message.&lt;&#x2F;p&gt;
&lt;p&gt;The thing that I really like about the Circle of Questions activity is that when you have a mixed group of people—some of whom tend to dominate the conversation while others tend to stay quiet—it gives everyone a chance to ask their questions and have their opinions heard without feeling like the most boisterous people have the only valuable input. We certainly have a mix like that in our case, and this activity worked wonderfully.&lt;&#x2F;p&gt;
&lt;p&gt;In the end, we came away with a number of things to change for the next release cycle. It&#x27;s interesting to note that none of our action items were directly related to the insights generated from the time line. I think the insights are still valuable (and are contained in the notes which we post to our internal wiki), so we may still come back to them in the near future; we were running over on time and had to wrap up even though the discussion could have continued longer.&lt;&#x2F;p&gt;
&lt;p&gt;Before we held this retrospective, I think some of the team members felt that we would not gain much from doing a full-day retrospective this time around. We had made good progress on our action items from the previous retrospective, and the general feeling was that things were going basically well. What could we possibly need to change?&lt;&#x2F;p&gt;
&lt;p&gt;People, please don&#x27;t hold retrospectives only when your team feels like there&#x27;s a problem that needs to be dealt with. If you hold retrospectives on a regular basis, whether you feel like you &quot;need&quot; to or not, you will uncover issues before the become problems. In our case, we managed to uncover a number of issues during our retrospective that could easily have grown into much larger problems if we didn&#x27;t bother to think about them until the effect was obvious.&lt;&#x2F;p&gt;
&lt;p&gt;If you do have obvious issues, then you might choose a different set of activities than what we did for this retrospective. The time line activity is great for discovering potential improvements even when everything seems to be going well, so I suggest giving it a try the next time your team is tempted to skip a retrospective because no one thinks there&#x27;s any huge problems.&lt;&#x2F;p&gt;
</description>
      </item>
    </channel>
</rss>
