Keystone Applied Intelligence
Applied AI engineering for enterprise and higher-consequence environments
Keystone is an independent engineering and R&D practice used to build, test, and evaluate concrete AI mechanisms across conversational workflows, authorization-aware retrieval, endpoint-compatible evaluation, observability, and bounded runtime-control research.
Core local operation has been demonstrated on customer-controlled infrastructure with local models and no external model API. This is an implementation-specific boundary, not a claim about every Keystone workload or deployment configuration.
Related projects, inspected independently
Keystone's projects are related engineering instruments and workloads. They share design concerns and selected mechanisms, but should not be interpreted as one fully composed or universally validated production runtime.
Conversational workflows with retrieval and escalation
The default served application uses one EngageOrchestrator with retrieval authorization, RAG, escalation, audit records, and OpenTelemetry GenAI tracing. A separate experimental path contains five specialist agents across four coordination phases and optional NATS JetStream integration; it is not the default or demonstrated production distributed execution.
Retained internal eval: keystone-engage/agent-v1 · 100/100 (70/70 core-regression, 25/25 architecture, 5/5 edge-case)
Public repo: keystone-engage
Authorization-aware retrieval
Query-time role, classification, and client predicates are designed to prevent unauthorized rows from becoming candidates for ranking or model context, with regression coverage for the tested isolation paths. Evidence-backed responses and citations are implemented without claiming production identity integration.
Evidence: cross-client isolation regression test (denial verified on both filtered and unfiltered paths). No dedicated Counsel evaluation baseline yet.
Public repo: keystone-counsel
Profile-based evaluation for compatible HTTP endpoints
Keystone Verify runs declarative cases against endpoints that satisfy a defined profile and contract, applies deterministic assertions, and records latency, structured results, and run metadata. Keystone Ledger separately retains selected internal passing, failing, adversarial, and remediation evidence.
Public repos: keystone-verify · keystone-ledger
Implementation boundary
Mechanisms are attributed by workload
Retrieval filters, task state, escalation, tracing, evaluation, and audit integrity appear in specific repositories with different scopes. Similar names and concerns do not establish a shared deployed substrate.
- › Engage: conversational workflow and escalation
- › Counsel: authorization-aware retrieval
- › Verify: endpoint-compatible evaluation
- › Ledger: retained internal evidence
- › Runtime Validity: bounded Track A research
Operational discipline for the LLM era
Enterprise contact-center systems developed mature operational patterns for escalation, validation, routing, logging, human handoff, and production recovery long before current LLM systems encountered analogous concerns. Keystone applies that experience to concrete AI mechanisms while keeping each implementation's evidence and limits explicit.
Workload-attributed controls, not universal claims
Some requirements are implemented outside the model in retrieval predicates, orchestration logic, database constraints, and audit mechanisms. The descriptions below identify where those mechanisms are demonstrated; they do not imply that every control operates in every workload.
Evidence-backed answers
Gov and relevant conversational paths return source-oriented answers and apply evidence or refusal rules at their evaluated boundaries.
Fail-closed refusal
Named Gov and Engage paths implement refusal behavior for tested insufficient-evidence or out-of-scope cases.
Query-time role-based access control
Counsel and relevant retrieval implementations apply query-time predicates before protected candidates enter ranking or model context. Regression tests cover the named isolation paths; implementation or configuration defects remain possible.
Human review for high-consequence actions
Engage implements severity-tier escalation and human handoff behavior in its conversational workflow.
Factual consistency scoring
Gov uses Vectara HHEM-2.1-Open as a model-based consistency mechanism. It is not an independent validator; scorer failure currently degrades open.
Audit trail with integrity checks
Engage and Counsel implement unkeyed SHA-256 hash chains; Gov uses a keyed HMAC-SHA256 value per record with limited signed-field coverage. These workload-specific mechanisms do not prove semantic correctness, authorization, or justification.
Selected claims tied to retained public artifacts
The public ledger retains selected internal quantitative baselines, remediation history, and failing-to-passing lineage. Each result belongs to its named workload, evaluated commit, configuration, and date. It is not independent validation.
- Precision@1 0.75 · MRR 0.79
- Corpus: 53 documents, 2,674 chunks (Alberta OHS)
- Hybrid retrieval: pgvector + full-text search
- Adversarial ACL: 8/8 blocked, 0 leaks
- Fail-closed: 5/6 (83%)
- Audit receipts logged (hash-chain fields empty in the retained dump; re-verification has not yet been recorded)
- Governed agent extension
- Corpus: 135 documents, 23,684 chunks
- 186 eval cases, 12 categories
- 558 executions, 0 failures
- 153 strict pass · 33 characterization
- 4 system bugs found by eval, all fixed & re-verified
- Enterprise HA or disaster recovery
- Multi-node or distributed deployment
- OIDC/SAML production identity integration
- Third-party penetration testing
- WCAG accessibility compliance
Recent work: governed-incident-agent, generative UI for governed agent actions, built at AI Tinkerers Global Hackathon, May 2026.
Built by Arnaldo Sepulveda
Keystone Applied Intelligence is the independent engineering practice of Arnaldo Sepulveda. The approach comes directly from 12+ years in enterprise contact-center systems at Genesys. These were environments that required production reliability, traceability, disciplined incident reconstruction, and careful coordination across system boundaries.
That work spanned Knowledge Center, classification systems, Digital Services, Agent Workspace, customer and interaction data, routing, conversational systems, integrations, migrations, go-lives, and production incident response. That operational experience informs Keystone's applied AI engineering without turning historical systems into claims about current LLM work.
- Practice Independent engineering practice of Arnaldo Sepulveda · based in Canada
- Background 12+ years at Genesys (enterprise contact-center and cloud systems)
- Scope Knowledge Center · classification · Digital Services · Agent Workspace · routing · integrations
- Domains Regulated and public-sector deployment experience
- More arnaldosepulveda.com