ADR-070: Multi-Domain RAG Architecture
Status
Proposed
Date
2025-12-27
Context
The current RAG implementation uses a single embedding space for all knowledge. As the system grows to support multiple domains (iDempiere core, CloudEmpiere, Angular, mobile, business, support, DevOps), we need domain isolation to:
- Avoid term collision - "component" means different things in Angular vs OSGi
- Enable precision - Queries return results from the correct context
- Support RBAC - Different user roles access different domains
- Scale independently - Update domains without affecting others
This ADR focuses specifically on domain separation and federated search. Related concerns are addressed in:
- ADR-071 - Documentation generation pipeline
- ADR-072 - Qute Web documentation server
- ADR-073 - Security, RBAC, rate limiting
Decision
1. Seven Knowledge Domains
| Domain | Content | Primary Sources |
|---|---|---|
idempiere |
Core ERP (AD, processes, callouts, validators) | Wiki, AD Metadata |
cloudempiere |
CloudEmpiere customizations, K_Entry articles | K_Entry, Git repos |
angular |
Angular frontend patterns, PrimeNG, SSR | Component docs |
mobile |
Mobile app architecture, offline sync | App docs, API specs |
business |
Business processes, workflows, SOPs | Process docs |
support |
Support KB, troubleshooting, FAQs | K_Entry, tickets |
devops |
Infrastructure, deployment, runbooks | Playbooks |
2. Single-Table with Domain Column
Instead of separate tables per domain, we use a single cli_embeddings table with a domain column:
-- Existing table, add domain column
ALTER TABLE cli_embeddings
ADD COLUMN IF NOT EXISTS domain VARCHAR(50) DEFAULT 'idempiere';
CREATE INDEX idx_cli_emb_domain ON cli_embeddings(domain);
CREATE INDEX idx_cli_emb_domain_source ON cli_embeddings(domain, source_type);
Why single table?
- pgvector handles large tables well with IVFFlat indexing
WHERE domain IN (...)with B-tree index is fast- Simpler migrations and maintenance
- No need for UNION views for cross-domain queries
3. Query Router
Pattern-based domain detection with AI fallback:
@ApplicationScoped
public class QueryRouter {
private static final Map<String, Pattern> DOMAIN_PATTERNS = Map.of(
"idempiere", Pattern.compile(
"(?i)\\b(AD_|C_Order|C_Invoice|callout|validator|process|window|tab|field)\\b"
),
"angular", Pattern.compile(
"(?i)\\b(component|directive|service|NgModule|Observable|PrimeNG|template)\\b"
),
// ... other domain patterns
);
public List<DomainWeight> detectDomains(String query) {
// Count pattern matches, normalize weights
// Return sorted list of domains with weights
}
}
Implementation Status:
- QueryRouter.java:22-266 - Fully implemented with 7 domain patterns
- Supports explicit domain prefix:
angular: how to create component
4. Federated Search Service
Search across multiple domains with weighted results:
@ApplicationScoped
public class FederatedSearchService {
public FederatedSearchResponse search(String query, int limit, List<String> domains) {
// 1. Detect domains via QueryRouter
// 2. Generate embedding for query
// 3. Single SQL query with domain IN clause
// 4. Return results with domain metadata
}
}
Implementation Status:
- FederatedSearchService.java:24-236 - Fully implemented
- Uses single SQL query with
WHERE domain IN (...) - Returns
FederatedSearchResponsewith domain stats
5. Domain Registry
Metadata table for domain configuration:
CREATE TABLE knowledge_domains (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
name VARCHAR(50) NOT NULL UNIQUE,
display_name VARCHAR(100) NOT NULL,
description TEXT,
is_active BOOLEAN DEFAULT true,
priority INTEGER DEFAULT 100,
required_roles TEXT[],
icon VARCHAR(50),
color VARCHAR(7),
created_at TIMESTAMP DEFAULT now()
);
Implementation Status
| Component | File | Status |
|---|---|---|
| Query Router | QueryRouter.java |
Implemented |
| Federated Search | FederatedSearchService.java |
Implemented |
| Domain Access | DomainAccessService.java |
Implemented |
| Domain Registry | DomainRepository.java |
Stub |
| Database Schema | V1__multi_domain_knowledge_schema.sql |
Created |
| MCP Tools | McpKnowledgeTools.java |
Wired |
MCP Tools Added
Three new domain-aware MCP tools were added to McpKnowledgeTools.java:
| Tool | Description |
|---|---|
searchDomain |
Federated search with domain parameter, auto-detection |
listDomains |
List available domains with document counts |
getDomainStats |
Get per-domain document statistics |
Consequences
Benefits
- Precision: Queries return results from correct context
- Isolation: Update domains independently
- Security: Per-domain RBAC via required_roles
- Performance: Smaller search space per query
Drawbacks
- Complexity: Need to maintain domain patterns
- Ingestion: Must tag content with correct domain
- Cross-domain: Some queries legitimately span domains
Next Steps
- ~~Run migration to add
domaincolumn tocli_embeddings~~ ✅ Done - ~~Wire
FederatedSearchServiceinto MCP tools~~ ✅ Done - ~~Create domain-specific MCP tools~~ ✅ Done (
searchDomain,listDomains,getDomainStats) - Add domain statistics endpoint:
/api/knowledge/domains - Implement
DomainRepositorywith database access - Ingest content into domains other than
idempiere