Insights
The hard part starts after the demo works.
Highlights
Briefings
Evals & quality
Make the answer checkable
Evals & quality
Close the loop from production to eval
Evals & quality
Trust the judge only after you check it against people
Integration & deployment
Pick how your agent connects to a client's systems
Conversation design & agentic
Design the conversation, not just the prompt
Conversation design & agentic
When a conversation is not enough
Conversation design & agentic
Evaluate the trajectory and the reliability
Conversation design & agentic
Voice changes the harness
Integration & deployment
Deploy inside the client's boundary
Personality & character
Sycophancy is measurable, so measure it
Personality & character
A guide that never disagrees is not a guide
Evals & quality
Choose and swap models against your own bar
Integration & deployment
The experience is made or lost in the handoffs
Safety & boundaries
Jailbreaks are a tax you keep paying
Safety & boundaries
Prompt injection is the attack you cannot fully patch
Safety & boundaries
Red-team the product before your users do
Trust, memory & governance
Keep the expert in the loop
Trust, memory & governance
What you are allowed to keep
Safety & boundaries
When not to answer, and when to bring in a person
Integration & deployment
Prove it to the security team
Personality & character
A prompt imitates a personality. A harness governs one.
Personality & character
Steer your assistant's personality on purpose
Trust, memory & governance
When remembering backfires
Evals & quality
Evals are the durable advantage for AI guidance products
Evals & quality
Turning a method into product behavior
Evals & quality
What to actually measure in conversational guidance
Evals & quality
Catch the regression before your users do
Trust, memory & governance
Shipping AI into governed environments
Trust, memory & governance