← All work

Product judgment under ambiguity

Harbour

A dashboard I built for my own real estate clients, run as a hypothesis test with the failure conditions written down before anyone logged in.

Role
Solo. Product, design and direction, with the code written by AI.
Timeline
Started Sept 7, 2026
Team
Me, with my broker and an attorney as domain reviewers
Status
Live in production with one pilot client
Repository
BT-Product/harbour-dashboard
Tools
Next.js, Supabase, Vercel, Claude Code
3 thresholds
Written before launch. None met yet.
9 decisions
Logged, including the ones that reversed me
1 client
Below my own usage threshold, and I'm not reading it yet

The short version

  • I’m a working agent, and I built this for my own clients. It shows a buyer or seller where their transaction stands, what happens next, and what I think about the homes they’ve seen.
  • It’s a hypothesis test, not a launch. I wrote the success thresholds and the failure condition before the first client logged in, so I can’t move them afterward.
  • Most of what I’ve learned came from being corrected. By a lender’s letter, by my own usage data, by my broker, and by an attorney. Four of the nine decisions in the log reversed something I believed.

Why this exists

Buying or selling a home is weeks of silence broken up by jargon. Clients don’t know what’s happening, so they call their agent, who repeats the same status update to every client that week.

The case I designed for is the move-up buyer, someone selling one home and buying another at the same time. They carry two transactions that have to land in the right order, and the question that keeps them awake is whether the sale closes in time to fund the purchase. A pure buyer has one timeline. A pure seller has one. Nobody builds for the hard case, so I did, and the simpler cases come along for free.

Harbour’s client overview: a plain-language status line, a wire fraud warning, and cards for upcoming tours, homes seen, and the purchase timeline. The client overview, from a demo account. One sentence of status, then a card for each section.

How I’m running it

Harbour is a hypothesis test with the conditions committed in writing before the first client saw it.

The hypothesis: if clients can see their transaction themselves, they’ll feel more looked after, and I’ll spend less time repeating status.

The riskiest part is “more looked after.” Self-serve status could just as easily replace the human contact that earns referrals. So the test treats heavy usage with a thinner relationship as a failure, not a win.

The thresholds, written before launch:

What I committed to Why it’s the bar
A median of 2+ visits a week per active client Habit, not a novelty click
No client saying at close that they felt less attended to The failure condition, stated as plainly as the success one
At least one client mentioning it unprompted Evidence they valued it enough to bring it up themselves

I wrote this into the strategy so success couldn’t be redefined after the fact. The repository is public, including the decision log and the day-by-day change log, so the reasoning can be checked rather than taken on my word. That matters more to me than the numbers themselves, because I’m both the builder and the agent in the room, and I’m the person most likely to talk myself into a good result.

The metric found a usability bug before it found a metric

Visit tracking shipped before the first client logged in, to measure retention. Its first useful output had nothing to do with retention.

My first client’s first session: six sections opened in 37 seconds, three to eight seconds each, then back to the start. That isn’t reading. That’s someone opening doors to see what’s behind them. The overview told them where they stood but nothing about what the menu contained, so clicking everything was the only way to find out.

I rebuilt the overview that week as a map: one plain sentence of status, then a card per section showing what’s inside. I could never have found this by using the product myself, because I already knew what was behind every door.

The principle: instrument early. Behavioral data earns its keep long before there’s enough of it for statistics.

Four times I was wrong

The lender owns the number. My first pre-approval screen rebuilt the loan math and displayed $492,228. The lender’s letter said $475,000. The arithmetic was right and the premise was wrong. The screen now shows the letter’s number verbatim and never recomputes it. The authoritative document sets the ceiling. The product may qualify it, never restate it.

We were telling clients to wire money with no warning. While preparing copy for broker review, I searched the codebase for “wire” and found exactly one line: an instruction to wire closing funds. No fraud warning anywhere, in a product for the single most targeted moment in a real estate transaction. It shipped that day, ahead of the review. Then practice corrected me again: escrow does email clients a portal link, so “we’ll never email you” was wrong. The warning now names the escrow officer and gives a number to verify against.

Harbour’s timeline view, with a wire fraud warning above the stage list for a purchase escrow. The warning that shipped the same day I noticed it was missing, with the wording corrected afterward by people who do escrow every week.

My broker killed a feature I’d asked for. I wanted buyers to see how a home’s HOA dues change what they can afford. I scoped it carefully so it could only ever lower the lender’s number, and shipped it. My broker ruled it out anyway: amortization is the lender’s licensed work, and an agent modeling it takes on liability that isn’t theirs. It’s being removed. Her reasoning also showed the way back, because anything the lender provides can be displayed. So I’ll ask my lender for adjusted figures at several HOA levels and show their table instead.

Harbour’s financials screen showing a pre-approval amount and an HOA affordability calculator. The screen as it looks today. The calculator below the pre-approval is the feature being removed, and the figures will come from the lender instead.

The product is a system of record, and I didn’t know it. The same review turned up something nobody had asked about. Every note an agent writes, on any platform, carries a duty to file it with their broker, who owns the file. Anything missing can weaken errors and omissions protection. That reframed the whole product. “Private” notes are private from the client, not the broker. The app has to export a handover file. And “remove client,” which deleted everything, destroyed exactly what was needed at exactly the moment it was due.

The principle across all four: in a licensed industry, what the product may say is part of what it is. Compliance is a design input, not a review at the end.

The best idea in the cycle wasn’t mine

My broker said the client’s stages shouldn’t be named after activities, like inspection, appraisal and loan approval, but after contingencies, so clients can see which protections they’ve released and which still stand.

She’s right, and it’s a better answer to the question a buyer in escrow is actually asking. An appraisal happening is an activity. A released appraisal contingency is their deposit at risk. The stage model is being rebuilt around the standard California contingency removal form.

The skill I’m practicing here isn’t having the idea. It’s noticing when the domain expert has just reframed the problem, and changing the product instead of defending the version I’d already built.

Where it stands, honestly

  • Started September 7, 2026. Live in production, with instrumentation shipped before the first client logged in.
  • One pilot client, onboarded September 10, who has opened it on two days. That’s below my own threshold, and a cohort of one can’t be read either way.
  • Not yet a move-up buyer, which is the case the hypothesis actually depends on.
  • Broker review complete. An attorney has given a direction, not yet answers.
  • Nothing about “launch” happens before three clients, including a move-up buyer, and a read against the thresholds.

How it gets built

I specify, review and decide. The code is written with AI, in Claude Code. I make the architecture calls, read what comes back, and debug when it breaks, but I’m not writing the backend by hand.

What that buys is speed at the part I’m good at: 19 database migrations and a live product in two weeks, multi-tenant from the first migration so a second agent is a new account rather than a rebuild. What it demands is judgment about what’s worth building, which is the work in every decision above.

What I’m carrying forward

Write the failure condition down before launch. “High usage with a thinner relationship is a failure” is the sentence that keeps this test honest, and I wrote it while I still had no data to be defensive about.

Ask practitioners what the product is, not just what it should do. The broker review didn’t produce a feature list. It told me the product was a system of record, which changed the data model.

Read your own instrumentation before you trust your own story. I assumed my pilot client used the dashboard daily. The data said two days, and every document I’d written had to be corrected.