Public-interest AI research case study
Building Observed: evidence-based organisational research, held to academic standard.
How Changeable designed Observed, a public-interest research platform that compares organisations against peer-reviewed academic benchmarks using only publicly available evidence, with a governance framework strict enough to support named publication.
Project overview
Public evidence. Fair analysis. Better accountability.
Most organisational accountability today relies on complaints, whistleblowers or investigative journalism, each valuable, but each dependent on someone being willing to come forward. Changeable built Observed to test a different approach: compare an organisation’s publicly available evidence, reviews, regulatory records, published reports, media coverage, against a peer-reviewed academic benchmark for what healthy organisational practice looks like.
The defining decision was to build the evidence base first. Before a single line of the analytical engine was written, an eight-domain academic benchmark framework was constructed from primary research literature, government guidance and regulatory sources, because the credibility of any finding depends entirely on the standard it is measured against.
Comparison, not accusation
The benchmark generates the standard. Observed measures organisations against it rather than making its own claims.
Public evidence only
No surveillance, no private data, no anonymous tip lines treated as fact. Every signal is a traceable public source.
Human review before publication
No output is published without a completed human review checklist and a named organisational right of response.
Suppression by design
Named findings require at least three independent source types before publication is even considered.
The problem
Public evidence about organisational practice already exists. Nobody structures it.
Regulatory findings, review platform patterns, published reports and media coverage about an organisation’s practices already sit in the public domain. What is missing is a rigorous, consistent way to read that evidence against an actual standard, rather than reacting to whichever single incident happens to attract attention.
A single bad review, one news story or an anonymous comment tells you very little on its own. The question Observed was built to answer is different: what does the full pattern of publicly available evidence say when measured against what credible research defines as healthy organisational practice?
Why this needed AI, and why it needed restraint
- Reading dozens of public sources against an academic framework by hand does not scale
- An AI system without discipline risks treating a single anonymous comment as fact
- The credibility of any finding depends on the rigour of the benchmark behind it
- Publication involving a named organisation carries real legal and reputational weight
- Governance had to be built in from the first design decision, not added later
The benchmark framework
Eight academic domains, built before the engine that uses them
Each domain framework was built by reading primary source documents directly, extracting definitional passages in the original authors’ own words rather than an AI-generated summary, so every claim traces back to a real, citable work.
Workplace bullying
Definitional standards and behavioural indicators drawn from established organisational psychology research and NZ regulatory guidance.
Psychosocial risk
Drawing on established psychosocial safety climate research and current NZ and Australian regulatory codes of practice.
Leadership and governance
Governance standards, director guidelines and leadership literature covering both healthy and toxic leadership patterns.
Psychological safety
Grounded in the foundational academic research defining team-level psychological safety and its measurement.
Worker wellbeing
Positive organisational research on thriving at work, alongside current NZ and Australian wellbeing survey data.
Impacts and consequences
Peer-reviewed evidence on the organisational and economic cost of poor practice, including return-on-investment research.
Legal and regulatory
Relevant NZ and Australian legislation and regulator guidance mapped directly against each other domain.
Cross-sector patterns
Comparative research across contracting, not-for-profit and high versus low-performing organisational contexts.
How it was built
The evidence base first, the engine second, the interface last
Observed was deliberately built in reverse order from a typical product build: the academic rigour came before a single feature of the analytical engine.
Decide what Observed is not, before deciding what it is
The founding decisions were as much about restraint as capability: public signals plus academic benchmarks plus gap analysis, never private data or surveillance. A minimum evidence threshold before any named finding could even be considered. A defined founding focus area to prove the method against before expanding scope. These constraints were locked before any research sourcing began, because a platform built to hold organisations to account has to hold itself to a higher standard first.
Read the primary research directly, one domain at a time
Rather than let an AI system synthesise a general understanding of workplace practice, each of the eight benchmark domains was built by reading academic and regulatory source documents directly and extracting genuine definitional passages in the original authors’ own words. This produced a two-layer knowledge architecture: structured framework summaries built around direct quotes for day-to-day use, with the full source documents held separately for verification. Every claim the engine makes can be traced back through that chain to a real, citable publication.
Document the process in full before automating any of it
Before the analytical engine was built, the entire operational process was documented: how a matter is triggered, how sources are collected and logged, how duplicate or corroborating signals are counted, and the exact sequence the analysis follows from raw evidence through to a structured, confidence-rated finding. This produced a full operator instruction set and a structured matter intake template, so the engine was built to match a proven process rather than the process being reverse-engineered from whatever the AI happened to produce.
Suppression rules, confidence ratings and right of response, built in from day one
A minimum source-diversity threshold, a defined confidence rating scale, a mandatory human review gate, and a formal right-of-response process were all built as core features of the methodology, not compliance add-ons applied after the fact. Legal review triggers were defined for the categories of finding that carry the greatest risk, and a conflict-of-interest policy was written to keep the platform independent of Changeable’s own client relationships. This is the same principle Changeable applies in client AI governance engagements: the controls have to be designed alongside the capability, not retrofitted once something has already gone wrong.
A public-interest evidence platform, not a complaints service
The website was built to be explicit about what Observed is and is not: not a complaints platform, not a legal service, and not a self-declared authority on truth. Core pages were built covering methodology, the research process, standards and evidence sources, and a formal legal, ethics and corrections page, alongside structured intake forms for both general enquiries and requests for analysis, each carrying explicit declarations about the public, evidence-based nature of any submission.
Governance
Built to a standard that could withstand scrutiny, not just produce output.
Publishing named findings about real organisations carries real weight. Observed’s human review gate exists to make sure nothing reaches publication without every one of these questions being answered first.
Evidence discipline
Not every public mention carries the same weight
Every source Observed draws on is classified and weighted, so a single anonymous comment is never treated with the same confidence as a formal regulatory finding.
What this produced
A methodology, not just a tool
The result of building the evidence base first is a platform that can explain and defend every finding it produces, not just generate one.
Traceable citation chains
Every finding can be traced from the published output back through the framework summary to the original source document.
Comparative, not accusatory language
A compliant language framework defines approved and prohibited framing for every category of finding.
Positive benchmarks included
The framework documents what healthy practice looks like, not only what poor practice looks like.
Structured multi-format outputs
Findings are designed to translate into a written report alongside accessible supporting formats for different audiences.
A defined correction process
Clarification, correction, update and withdrawal are all named outcomes with their own documented process.
A platform, not a single report
The methodology, benchmark library and governance framework are reusable across every future matter, not rebuilt each time.
Questions
Questions about how Observed was built?
Common questions about the approach behind Observed’s evidence-based research methodology.
Why was the benchmark library built before the analytical engine?
Because the credibility of any finding depends entirely on the standard it is measured against. Building the academic benchmark first meant the engine was constrained by real research from the outset, rather than the framework being fitted around whatever the AI produced.
What sources does Observed use?
Only publicly available material: formal regulatory and legal records, parliamentary and official information material, media reporting, organisation-owned publications, public registers and public review platforms. No private data, surveillance or confidential information is used.
How does Observed avoid treating rumour as fact?
Every source is classified by type and weighted accordingly, a single anonymous review is never sufficient on its own, and named findings require a minimum of three independent source types before publication is even considered.
Does an organisation get a chance to respond before publication?
Yes. A formal right-of-response process is built into the methodology, with a standard 10 business day response window before any named finding proceeds to publication.
Is every finding automated by AI?
No. AI accelerates the evidence review and drafting, but every output passes through a documented human review checklist, and named findings above certain risk thresholds require legal review before publication.
How does Observed handle conflicts of interest?
A formal conflict register assesses every matter before analysis begins, and Observed will not publish findings about any organisation with a recent commercial relationship to Changeable or its associated businesses unless strict conditions are met.
What happens if a finding turns out to be wrong?
A defined correction process covers clarification, correction, update and withdrawal, each with its own documented process and a commitment to transparent, dated acknowledgement.
Want to see how Observed’s methodology works?
Explore Observed directly, or talk to Changeable about applying the same evidence-first, governed approach to your own AI product or research process.