Keystone Applied Intelligence

Applied AI engineering for enterprise and higher-consequence environments

Keystone is an independent engineering and R&D practice used to build, test, and evaluate concrete AI mechanisms across conversational workflows, authorization-aware retrieval, endpoint-compatible evaluation, observability, and bounded runtime-control research.

Core local operation has been demonstrated on customer-controlled infrastructure with local models and no external model API. This is an implementation-specific boundary, not a claim about every Keystone workload or deployment configuration.

Demo access demo.getkeystone.ai · operator1 / demo123 · try: “What atmospheric testing is required before entering a confined space?”
01 Engineering workloads

Related projects, inspected independently

Keystone's projects are related engineering instruments and workloads. They share design concerns and selected mechanisms, but should not be interpreted as one fully composed or universally validated production runtime.

Keystone Engage Conversational

Conversational workflows with retrieval and escalation

The default served application uses one EngageOrchestrator with retrieval authorization, RAG, escalation, audit records, and OpenTelemetry GenAI tracing. A separate experimental path contains five specialist agents across four coordination phases and optional NATS JetStream integration; it is not the default or demonstrated production distributed execution.

Retained internal eval: keystone-engage/agent-v1 · 100/100 (70/70 core-regression, 25/25 architecture, 5/5 edge-case)

Public repo: keystone-engage

Keystone Counsel Retrieval

Authorization-aware retrieval

Query-time role, classification, and client predicates are designed to prevent unauthorized rows from becoming candidates for ranking or model context, with regression coverage for the tested isolation paths. Evidence-backed responses and citations are implemented without claiming production identity integration.

Evidence: cross-client isolation regression test (denial verified on both filtered and unfiltered paths). No dedicated Counsel evaluation baseline yet.

Public repo: keystone-counsel

Keystone Verify Evaluation

Profile-based evaluation for compatible HTTP endpoints

Keystone Verify runs declarative cases against endpoints that satisfy a defined profile and contract, applies deterministic assertions, and records latency, structured results, and run metadata. Keystone Ledger separately retains selected internal passing, failing, adversarial, and remediation evidence.

Public repos: keystone-verify · keystone-ledger

Implementation boundary

Mechanisms are attributed by workload

Retrieval filters, task state, escalation, tracing, evaluation, and audit integrity appear in specific repositories with different scopes. Similar names and concerns do not establish a shared deployed substrate.

  • Engage: conversational workflow and escalation
  • Counsel: authorization-aware retrieval
  • Verify: endpoint-compatible evaluation
  • Ledger: retained internal evidence
  • Runtime Validity: bounded Track A research
02 Why this exists

Operational discipline for the LLM era

Enterprise contact-center systems developed mature operational patterns for escalation, validation, routing, logging, human handoff, and production recovery long before current LLM systems encountered analogous concerns. Keystone applies that experience to concrete AI mechanisms while keeping each implementation's evidence and limits explicit.

Contact centers solved
What LLM products rediscovered
Severity-tier escalation
Human-in-the-loop for high-consequence actions
Per-step validation
Evidence gating between agent steps
Compliance logging
Integrity-checked audit trails
Confidence-threshold refusal
Fail-closed when evidence is insufficient
Policy-aware routing
Authorization-first retrieval and dispatch
03 Controls

Workload-attributed controls, not universal claims

Some requirements are implemented outside the model in retrieval predicates, orchestration logic, database constraints, and audit mechanisms. The descriptions below identify where those mechanisms are demonstrated; they do not imply that every control operates in every workload.

Evidence-backed answers

Gov and relevant conversational paths return source-oriented answers and apply evidence or refusal rules at their evaluated boundaries.

Retrieval flow

Fail-closed refusal

Named Gov and Engage paths implement refusal behavior for tested insufficient-evidence or out-of-scope cases.

Retrieval flow

Query-time role-based access control

Counsel and relevant retrieval implementations apply query-time predicates before protected candidates enter ranking or model context. Regression tests cover the named isolation paths; implementation or configuration defects remain possible.

Database model

Human review for high-consequence actions

Engage implements severity-tier escalation and human handoff behavior in its conversational workflow.

Orchestration

Factual consistency scoring

Gov uses Vectara HHEM-2.1-Open as a model-based consistency mechanism. It is not an independent validator; scorer failure currently degrades open.

Orchestration

Audit trail with integrity checks

Engage and Counsel implement unkeyed SHA-256 hash chains; Gov uses a keyed HMAC-SHA256 value per record with limited signed-field coverage. These workload-specific mechanisms do not prove semantic correctness, authorization, or justification.

Audit chain
04 Evaluation Evidence

Selected claims tied to retained public artifacts

The public ledger retains selected internal quantitative baselines, remediation history, and failing-to-passing lineage. Each result belongs to its named workload, evaluated commit, configuration, and date. It is not independent validation.

keystone-core/retrieval-v1 · 2026-04-11 · published in keystone-ledger
  • Precision@1 0.75 · MRR 0.79
  • Corpus: 53 documents, 2,674 chunks (Alberta OHS)
  • Hybrid retrieval: pgvector + full-text search
  • Adversarial ACL: 8/8 blocked, 0 leaks
  • Fail-closed: 5/6 (83%)
  • Audit receipts logged (hash-chain fields empty in the retained dump; re-verification has not yet been recorded)
keystone-core/agent-v1 · 2026-05-20 · published in keystone-ledger
  • Governed agent extension
  • Corpus: 135 documents, 23,684 chunks
  • 186 eval cases, 12 categories
  • 558 executions, 0 failures
  • 153 strict pass · 33 characterization
  • 4 system bugs found by eval, all fixed & re-verified
Remediation · failing-to-passing lineage FC-005 domain scope guard, demo-grade, merged 2026-05-17. A pre-retrieval gate that refuses out-of-corpus queries (emissions regulations, workers comp, federal tax, IT procurement). It targets the retrieval-v1 FC-005 failure, where a greenhouse-gas query returned Part 36 mine-gas chunks via embedding overlap. Re-verification has not yet been recorded. The full taxonomy-based two-stage domain gate is not yet implemented; this is a partial, documented remediation. Separately, the failing keystone-core/agent-v0 run is retained alongside its passing agent-v1 remediation run. See the eval ledger.
Not claimed
  • Enterprise HA or disaster recovery
  • Multi-node or distributed deployment
  • OIDC/SAML production identity integration
  • Third-party penetration testing
  • WCAG accessibility compliance

Recent work: governed-incident-agent, generative UI for governed agent actions, built at AI Tinkerers Global Hackathon, May 2026.

05 About

Built by Arnaldo Sepulveda

Keystone Applied Intelligence is the independent engineering practice of Arnaldo Sepulveda. The approach comes directly from 12+ years in enterprise contact-center systems at Genesys. These were environments that required production reliability, traceability, disciplined incident reconstruction, and careful coordination across system boundaries.

That work spanned Knowledge Center, classification systems, Digital Services, Agent Workspace, customer and interaction data, routing, conversational systems, integrations, migrations, go-lives, and production incident response. That operational experience informs Keystone's applied AI engineering without turning historical systems into claims about current LLM work.

  • Practice Independent engineering practice of Arnaldo Sepulveda · based in Canada
  • Background 12+ years at Genesys (enterprise contact-center and cloud systems)
  • Scope Knowledge Center · classification · Digital Services · Agent Workspace · routing · integrations
  • Domains Regulated and public-sector deployment experience
  • More arnaldosepulveda.com