Some AI risks only become visible across a conversation rather than in a single response.
Dependency reinforcement, harmful validation, boundary confusion, poor crisis response, and emotional escalation can emerge over multiple turns.
How can teams test these behaviors systematically before they reach real users?
Who this is for
The initial audience is product, trust & safety, or compliance leads at small-to-mid-size companion and wellness AI companies. Before or during release decisions, they need to understand what failed, where it failed, and whether there is enough evidence to act.
This is the intended user hypothesis; external user validation is still needed.
What I built
I built an open-source evaluation harness that runs versioned scenarios against a target chatbot, captures the resulting multi-turn conversation, and evaluates it using a psychosocial safety rubric.
The broader product is designed around psychosocial safety evaluation. The current experimental v1 starts with one construct: relational sycophancy.
Relational sycophancy is when a chatbot gives unsupported weight to a user’s interpretation of another person or relationship—for example, affirming an uncertain assumption about someone’s motives as though it were established.
Each finding includes category, severity, exact evidence, rationale, and mechanisms.
Scenario pack
Persona
Intent
Risk context
Turn count
→
Multi-turn conversation
Simulated user
Follows scenario and user intent
⇄
Target chatbot
Model under evaluation
→
Evaluation judge
Applies rubric and identifies risks
→
Findings & report
Severity Evidence Rationale Mechanisms
Three key product decisions
1
Don’t reduce safety to one number
A single safety number would make models easy to compare, but it would hide what actually failed. A crisis-response failure and dependency reinforcement are not interchangeable.
Decision
Report construct-specific severity and attach exact evidence and mechanisms to every positive finding.
Tradeoff
Less convenient for ranking, much more useful for diagnosis and fixing.
2
Use repeatable simulations
Real-user transcripts are harder to curate into controlled tests of specific risk contexts. Versioned synthetic scenarios let me deliberately test those contexts under consistent conditions.
Decision
Use versioned scenario packs with personas, intents, risk contexts, and turn counts.
Tradeoff
Less real-world fidelity, but greater control over scenario coverage and test conditions.
3
Start with one construct first
Psychosocial safety spans several behaviors. A first release needs a clear definition of what it is evaluating.
Decision
Focus the experimental v1 on relational sycophancy, with a dedicated rubric and 20-scenario pack.
Tradeoff
Narrower coverage, but a more testable rubric, scenario pack, and validation target.
Scope
In v1
✓Relational sycophancy evaluation
✓English
✓Text-only conversations
✓Simulated users
✓Multi-turn scenarios
✓Category-level findings
✓Evidence and rationale
✓Human- and machine-readable outputs
Not in v1
×Other psychosocial constructs
×Certification
×Clinical diagnosis
×Real customer transcripts
×Full governance workflows
×Long-horizon agents
×Universal safety score
Validation approach
Implemented checks
•Structured output and schema checks
•Exact evidence checks against assistant turns
•Deterministic severity aggregation
Development review
The 20-scenario reference run was manually reviewed during development. That review identified one likely false negative, RS-004.
Planned independent validation
The judge is not yet formally validated against independent human annotations. The next validation step is independently annotated transcripts, blinded comparison, and explicit disagreement adjudication.
ⓘImplemented checks make results auditable; they do not establish judge validity.
As the architecture took shape, I realized that supporting more providers would not make the product more useful if teams could not trust the evaluation itself.
Implication
Evaluation quality—not provider breadth—was the more important product risk.
Changed priority
I shifted attention toward scenario quality, judge validation, and evidence-rich reporting.
Deferred
Broader provider integrations and additional architecture work.
Next test
Determine how reliably the evaluator agrees with human review and how stable findings are across repeated assessments.
The main product risk wasn’t whether the tool could connect to enough models. It was whether teams could trust the findings enough to act on them.
Current status
Experimental v1 released · validation ongoing. The current release evaluates relational sycophancy; independent human validation is pending.