Skip to content

Insights

The hard part starts after the demo works.

Intelligence is getting cheap. What lasts is the specialized system around the model, and the discipline to keep it improving as the frontier moves.

Newsletter

Reflective Surfaces

Our editorial publication on what makes a conversation good, the human craft as much as the technical work.

Read on Substack →

Highlights

Where we land on the questions teams ask us most.

Briefings

The disciplines behind a guidance product users trust, from defining what good looks like to holding the line in the hard moments.

29 articles

Evals & quality

Make the answer checkable

In guidance, a fluent answer that sounds right is not the same as one you can trace to a source. Close that gap on purpose.

July 1, 2026·7 min read
Evals & quality

Close the loop from production to eval

The product that improves every week is the one where each real-world failure becomes a permanent test the next release has to pass.

July 1, 2026·8 min read
Evals & quality

Trust the judge only after you check it against people

The model grading your answers has measurable biases toward the first option, the longer answer, and its own writing. Validate it before you trust it.

July 1, 2026·7 min read
Integration & deployment

Pick how your agent connects to a client's systems

Direct API, MCP, a connector, or an embed. The connection you choose sets how much you build, how much control you keep, and how much a security team has to review.

June 30, 2026·7 min read
Conversation design & agentic

Design the conversation, not just the prompt

A guidance product is a designed path through the moments that decide it, not one clever prompt. Map the states, the moves, and the transitions.

June 30, 2026·8 min read
Conversation design & agentic

When a conversation is not enough

Sometimes a guidance agent has to act, not only advise. Giving it real tools turns talk into an outcome, and every tool is a new door in.

June 30, 2026·8 min read
Conversation design & agentic

Evaluate the trajectory and the reliability

Once an agent takes actions with tools, grading the final answer hides where it went wrong and how often it will.

June 30, 2026·7 min read
Conversation design & agentic

Voice changes the harness

A voice agent is not a text agent with a speaker attached. It is a real-time pipeline where a lot of the character lives in timing.

June 30, 2026·8 min read
Integration & deployment

Deploy inside the client's boundary

Residency, gateways, and egress decide whether a conversational tool can ship into an enterprise at all. Design for them from the first sprint or they become a rebuild.

June 29, 2026·7 min read
Personality & character

Sycophancy is measurable, so measure it

The default an approval trained model hands you is a flatterer. For a coach or a tutor that is disqualifying, and you can put a number on it.

June 29, 2026·8 min read
Personality & character

A guide that never disagrees is not a guide

In coaching and care, the value is often the pushback. Calibrated disagreement is a behavior you design and grade, not one the model gives you for free.

June 29, 2026·8 min read
Evals & quality

Choose and swap models against your own bar

A new model ships every few weeks. Picking one and moving to the next is a re-score on your eval suite, not a leaderboard read.

June 29, 2026·7 min read
Integration & deployment

The experience is made or lost in the handoffs

Latency, streaming, and the seams where control passes decide how a deployed tool feels. Design the handoffs, or they surprise you in production.

June 28, 2026·8 min read
Safety & boundaries

Jailbreaks are a tax you keep paying

A safety boundary is not a bug you close once. Adversarial users keep finding the way around it, so robustness is a posture, not a milestone.

June 28, 2026·8 min read
Safety & boundaries

Prompt injection is the attack you cannot fully patch

When an agent reads a web page, a document, or a stored memory, that content can carry instructions it follows. It is a structural problem, not a bug.

June 28, 2026·8 min read
Safety & boundaries

Red-team the product before your users do

Adversarial testing is a standing discipline, not a one-time audit. For a guidance product, aim it at the trust moments.

June 28, 2026·7 min read
Trust, memory & governance

Keep the expert in the loop

Expert review is a quality lever and a compliance control at once, and in guidance products it is the same person doing both jobs.

June 28, 2026·7 min read
Trust, memory & governance

What you are allowed to keep

The sharpest privacy decision in a guidance product is not what it says back, it is what you store and what you train on.

June 28, 2026·7 min read
Safety & boundaries

When not to answer, and when to bring in a person

A block and a disclaimer are not a safety design. Refusal and escalation are features you design and test like any other behavior.

June 27, 2026·6 min read
Integration & deployment

Prove it to the security team

Logging, audit, and compliance evidence are what turn a working integration into one a client will actually deploy. Build the proof as you build the tool.

June 26, 2026·7 min read
Personality & character

A prompt imitates a personality. A harness governs one.

A system prompt sets the tone in a demo. Holding that character once real users arrive is an engineering problem.

June 26, 2026·9 min read
Personality & character

Steer your assistant's personality on purpose

Leave a conversational tool's character to chance and the training picks it for you, usually a people pleaser.

June 24, 2026·5 min read
Trust, memory & governance

When remembering backfires

More memory is not more helpful. Unscoped recall can make a product feel like it is watching you.

June 18, 2026·5 min read
Evals & quality

Evals are the durable advantage for AI guidance products

Model access is a commodity. The advantage that compounds is a domain eval suite only you can run.

May 28, 2025·8 min read
Evals & quality

Turning a method into product behavior

Coaching, teaching, and care methods are full of judgment that lives in one person's head. Here is how that judgment becomes something you can build and grade.

May 14, 2025·8 min read
Evals & quality

What to actually measure in conversational guidance

Vanity metrics won't tell you if the product helped. Measure whether the system followed the method, turn by turn.

April 30, 2025·7 min read
Evals & quality

Catch the regression before your users do

A prompt tweak that helps one case can quietly break ten others. Gate every change on the full suite and prove the win with error bars.

April 16, 2025·8 min read
Trust, memory & governance

Shipping AI into governed environments

Privacy, residency, and incoming regulation aren't afterthoughts. Read the obligations by name and design to them from the first sprint.

March 26, 2025·8 min read
Trust, memory & governance

Memory is the trust surface

The moment a system remembers you, it takes on a responsibility it didn't have before.

March 12, 2025·6 min read