Skip to content

Under active development

Development status →
TACITUS
Section pages

Research experiments · revised 29 September 2026

Questions we can put to a test.

Structured knowledge is useful only if it improves the work. These proposed studies examine evidence selection, total cost, error detection, and continuity between people.

Under development. These are study designs. We are not reporting completed experiments, recruitment, measured savings, or comparative performance. PRAXIS remains the main product focus.

PROPOSED STUDY 01

Does a smaller context preserve the decisive evidence?

Comparison
Use a fixed dossier with questions requiring qualifications, cross-document links, and contrary accounts. Compare full-document context, tuned retrieval, retrieval plus compression, and capsule assembly with the same model and answer instructions.
What to measure
Evidence coverage, supported answers, missing exceptions, and citation accuracy at several input budgets. Include extraction, indexing, updates, retries, latency, and reviewer minutes in a separate total-cost comparison.
What would count against the design
Reject an efficiency claim if fewer tokens produce more material omissions, or if preparation and correction cost more than they save over the stated workload.
Context assembly and baselines

PROPOSED STUDY 02

Can the next analyst reconstruct the reasoning?

Comparison
Give independent reviewers the same case through a conventional briefing, a briefing with linked sources, or a capsule with review records. Ask them to explain a prior decision, identify a planted error, and update the assessment.
What to measure
Time to locate the relevant passage, missed caveats, incorrect approvals, and quality of the revised artifact. Blind output review to the condition and preserve disagreements between reviewers.
What would count against the design
The structure has not earned its cost if reviewers merely feel more confident while missing the same errors or spending longer reconstructing the case.
What a reviewer needs to check

PROPOSED STUDY 03

Does a correction reach every dependent answer?

Comparison
Run questions over a versioned case, then add a source correction, withdraw an interpretation, or change access permissions. Repeat current-state and historical questions separately.
What to measure
Stale answers, invalidated summaries, incorrect hindsight, unauthorized retrieval, and the cost of updating derived records. Save both the source snapshot and the context actually used.
What would count against the design
A reusable record fails if it keeps serving an obsolete premise without warning or rewrites what was knowable before the correction.
Time, corrections, and historical questions

Scenario and dialogue research

Wind Tunnel and CONCORDIA extend the same questions into institutional work. A plausible scenario is not a prediction; a tidy transcript is not an agreed account.

Wind Tunnel

Compare plausible responses to a policy proposal while preserving the assumptions behind each scenario. A proposed evaluation would ask whether independent analysts identify missing constraints and alternative explanations, using an ordinary scenario workshop as the baseline.

Generated reactions must remain hypothetical. This research supports analysis and preparation; it does not establish predictive accuracy or optimize manipulation.

Experimental project ↗

CONCORDIA

Organize a dialogue into attributed statements, disputed interpretations, and possible commitments. Compare the result with ordinary minutes on missed qualifications, attribution errors, and participants’ ability to correct the record.

Consent, confidentiality, and participant review belong in the study design. An inferred interest or proposed commitment must not be recorded as something a participant accepted.

Experimental project ↗

What a published result should include

Freeze the questions, sources, model versions, and scoring rubric before the held-out run. Publish raw outputs where permitted, failures and retries, paired uncertainty estimates, reviewer disagreements, and the cost assumptions. Include a strong baseline and explain which findings may transfer beyond the tested task.

The approach draws on NIST’s ARIA evaluation planning manual, which combines model, adversarial, and user testing. The study designs and failure criteria above are our proposals.

Search and terminologyMachine-readable index