Watch a recommendation become an auditable decision.
Choose a synthetic policy scenario or let the Auto Tour cycle through them. The left side shows what a shallow policy-matching approach might conclude. The right side shows the public VERITAS workflow: ground relevant rules, identify normative constraints, check evidence and exceptions, resolve conflict or precedence, then produce an explicit decision and trace.
The recommendation must be checked against applicable policy obligations and available evidence before execution.
Recommendation appears consistent with policy.
Matched authorization language and found no direct prohibition.
Action satisfies the applicable policy constraints.
Required authorization is verified and the decision can be logged with sufficient evidence.
Evaluate the hard cases: exceptions, conflicts, missing evidence, and policy precedence.
The benchmark contains 800 deterministic synthetic policy-decision cases. Four hundred cases are used only to calibrate confidence. A disjoint 400-case set is used for every headline result below. Cases span eight classes and explicitly include situations that shallow keyword or rule matching often mishandles.
overall decision accuracy
Correct ALLOW, BLOCK, or ESCALATE outcome on held-out cases.
violation-detection F1
Precision and recall for cases whose correct outcome is BLOCK.
exception-handling accuracy
Valid and invalid exception scenarios evaluated correctly.
conflict / knowledge-limit escalation
Cases requiring human review or explicit escalation are recognized.
complete decision traces
Required synthetic policy, evidence, decision, and reasoning fields present.
Same 400 cases, transparent baseline vs. public VERITAS method.
| Metric | Shallow baseline | VERITAS demo |
|---|---|---|
| Overall decision accuracy | 53.8% | 93.8% |
| Violation precision | 62.1% | 95.5% |
| Violation recall | 65.5% | 95.0% |
| Violation F1 | 63.7% | 95.2% |
| False-block rate on allowable actions | 28.0% | 5.0% |
| Exception-handling accuracy | 51.0% | 94.0% |
| Escalation accuracy | 16.0% | 91.0% |
| Audit-trace completeness | 41.3% | 95.8% |
| Cause / evidence alignment | 53.7% | 97.9% |
| Expected Calibration Error | 5.6% | 1.4% |
Policy logic beyond direct prohibitions
- Compliant action: all obligations and evidence requirements satisfied.
- Direct prohibition: an applicable MUST NOT rule blocks the action.
- Missing prerequisite: required authorization or evidence is absent.
- Valid exception: an exception overrides the general prohibition because all exception conditions are satisfied.
- Invalid exception: the exception is invoked but required conditions or evidence are missing.
- Rule conflict: applicable policies point to incompatible outcomes and require escalation.
- Policy precedence: a newer or higher-priority rule changes the applicable outcome.
- Insufficient evidence: the system should not force ALLOW or BLOCK when knowledge limits are reached.
Policy to constraint to evidence to decision to trace.
The public architecture shows functional responsibilities only. It intentionally does not disclose proprietary semantic-grounding methods, normative compilation, conflict-resolution logic, optimization procedures, model configurations, or implementation-specific scoring.
Decision context
Receive the proposed action, relevant facts, available evidence, actors, resources, and operating context.
Policy relevance
Identify potentially governing provisions and connect policy concepts to the decision context.
Normative constraints
Represent obligations, prohibitions, permissions, exceptions, evidence conditions, and applicable versions.
Candidate decisions
Test the proposed action and feasible alternatives against the grounded constraints and evidence state.
Conflict & knowledge limits
Handle exceptions, precedence, incompatible rules, missing evidence, and cases requiring human review.
Decision evidence
Return the decision, grounded rules, supporting evidence, reasons, confidence, alternatives, and auditable trace.
The demo is designed around measurable decision-assurance questions.
Did the system ground the right policy?
Measure whether the governing provisions for a synthetic decision are identified and carried into the decision trace.
Are violations caught without excessive blocking?
Track violation precision/recall together with the false-block rate on otherwise allowable actions.
Are exceptions handled correctly?
Distinguish valid exceptions from incomplete or unsupported attempts to invoke an exception.
Does the system recognize knowledge limits?
Escalate conflicts and insufficient-evidence cases rather than fabricating a confident binary outcome.
Is the decision auditable?
Require matched policy, evidence state, decision, reason, confidence, and trace metadata for review.
Does confidence correspond to correctness?
Calibrate confidence on one split and measure Expected Calibration Error only on disjoint held-out cases.
Show the evidence. Protect the decision-assurance recipe.
What evaluators can inspect
- Problem definition and decision-assurance workflow.
- Generic policy concepts: obligations, prohibitions, permissions, exceptions, precedence, and evidence requirements.
- Synthetic benchmark structure, held-out split, metrics, and limitations.
- Illustrative decision records and audit fields.
- Transparent shallow baseline used for comparison.
Protected implementation details
- Proprietary semantic-grounding or policy-retrieval methods.
- Internal normative-logic compilation and representation.
- Detailed conflict-resolution and candidate-selection procedures.
- Private prompts, model configurations, thresholds, or scoring functions.
- Source code, private datasets, customer policy documents, or protected integrations.
Policy-grounded decision assurance is broader than any single industry.
AI Governance & Risk
Check AI-enabled recommendations against organizational policies, evidence requirements, and defined escalation boundaries.
Compliance Operations
Connect operational decisions to explicit rules, exceptions, approvals, and reviewable evidence without relying on opaque reasoning alone.
Financial & Enterprise Controls
Support approval, access, transaction, and model-governance decisions with traceable policy checks and evidence gating.
Healthcare Administration
Apply policy-grounded checks to administrative workflows where permissions, prerequisites, documentation, and human review matter.
Industrial & Safety-Critical Workflows
Make constraints, required evidence, exceptions, and escalation conditions visible before an automated recommendation becomes an action.
Human-AI Decision Systems
Provide reviewers with grounded reasons, alternatives, uncertainty, and audit evidence instead of a recommendation alone.
Systems, formalized decision assurance, and AI/ML in one research team.
Dr. Sajib Datta
Technical direction, policy-to-system architecture, evidence and audit pipelines, experimental design, reproducibility, performance evaluation, commercialization-oriented prototyping, and end-to-end integration.
Dr. Tonmoay Deb
Semantic reasoning, trustworthy AI, confidence-aware decision methods, adversarial and edge-case analysis, model evaluation, explanation design, and AI/ML experimentation.
Grounded in explainability, knowledge limits, risk measurement, and documented decision processes.
The AI RMF organizes risk-management activity around Govern, Map, Measure, and Manage and emphasizes continuous risk management across the AI lifecycle.
Public reference ↗NIST identifies explanation, meaningfulness, explanation accuracy, and knowledge limits as fundamental principles for explainable AI systems.
Public reference ↗NIST's cross-sector GenAI profile extends AI RMF guidance for managing risks associated with generative AI systems across the lifecycle.
Public reference ↗VERITAS on this page is an independent Omniscient Innovations LLC research demonstration using deterministic synthetic policies, evidence records, and decision cases. It is agency-neutral and is not sponsored, funded, endorsed, certified, selected, or validated by any government entity, standards body, regulator, or other organization. The public workflow intentionally uses simplified decision-assurance logic and does not disclose proprietary algorithms, formalization methods, optimization procedures, private prompts, source code, customer information, controlled technical information, or confidential policy data. The benchmark does not establish legal or regulatory compliance, is not legal advice, and must not be interpreted as fielded, certified, independently validated, or production performance.