Back to all articlesSystem Integration

Tenant Context Is an Authorization Primitive: Isolation for Multi-Tenant AI Agents

Tenant IDs cannot live in prompts alone. Build authorization-first retrieval, isolated memory and caches, server-side tool gating, and deterministic row-level enforcement.

Mapki

Mapki

Oct 10, 2026•9 min read
Share:
Tenant Context Is an Authorization Primitive: Isolation for Multi-Tenant AI Agents

Tenant Context Is an Authorization Primitive: Isolation for Multi-Tenant AI Agents

A tenant ID is not a prompt variable.

If a multi-tenant AI agent receives tenant_id as ordinary text, then the model, user, retrieved document, or tool output can potentially influence it. That is not an authorization boundary. It is a suggestion.

In a production SaaS agent, tenant context must be established from authenticated identity and carried through every security-sensitive operation: retrieval, memory, caching, tool calls, queues, files, analytics, and final response assembly.

The central rule is simple:

Relevance decides which authorized data is useful. Authorization decides which data is eligible.

A vector search result that is highly relevant but belongs to another tenant must never become a candidate for ranking or generation.

Why multi-tenant agents fail differently

Traditional SaaS authorization usually protects an API request and a database query. Agents introduce more intermediate state and more ways to accidentally widen scope.

  • Retrieval: A shared vector index can return another tenant’s chunk before application code filters it.
  • Memory: A long-term memory store can associate facts with a user or tenant incorrectly.
  • Conversation state: A shared thread or cache can reuse context across sessions.
  • Tools: The model may choose a tool that operates under a broader service identity than the requesting user.
  • Queues: A background job can lose tenant context between enqueue and execution.
  • Logs and traces: Prompts, retrieved documents, and tool results can leak through observability systems.
  • Summaries: A compacted context or handoff can preserve facts while dropping the scope that made them legal to use.

The 2026 paper Securing the Agent formalizes the core gap: retrieval ranks by relevance, not authorization. It identifies tool-mediated disclosure, context accumulation, and client-side orchestration bypass as additional multitenant risks. Its layered architecture combines policy-aware ingestion, retrieval-time gating, and server-side orchestration, and reports that ABAC gating eliminated cross-tenant leakage in its open-source implementation with negligible overhead.

The reference architecture

Keep a trusted request context outside the model’s control.

Authenticated request
|
v
Identity service
|  verified subject, tenant, roles, scopes, policy version
v
Agent gateway
|  creates immutable request context
v
Server-side orchestrator
|  authorization-first retrieval + tool policy
+--------------------+---------------------+
v                    v                     v
Memory API         Retrieval API         Tool gateway
|                    |                     |
tenant-scoped      row/security trim     user-bound capability
stores             before ranking        and argument checks
+--------------------+---------------------+
v
Model endpoint
receives only authorized context
  • Identity service: Authenticates the caller and maps the subject to one or more allowed tenants.
  • Agent gateway: Creates a signed, immutable request context. The model cannot modify it.
  • Server orchestrator: Owns state transitions, tool routing, and authorization checks.
  • Data APIs: Encapsulate tenant and user filtering so agents cannot query stores directly.
  • Model: Reasons over the scoped context but does not define the scope.

Google’s multi-tenant agent reference architecture follows this pattern by extracting tenant identity at the frontend, applying policy boundaries around tenant resources, and propagating identity to shared MCP servers. Google recommends local MCP deployments for highly sensitive data and shared MCP servers for common corporate tools, provided user identity is securely propagated and enforced downstream.

Make request context tamper-resistant

The agent should receive a server-generated context handle, not authority fields it can rewrite.

interface RequestContext {
requestId: string;
subject: string;
tenantId: string;
roles: string[];
scopes: string[];
policyVersion: string;
issuedAt: string;
expiresAt: string;
signature: string;
}

function validateContext(ctx: RequestContext, token: VerifiedToken) {
if (ctx.subject !== token.sub) throw new Error("subject mismatch");
if (!token.tenants.includes(ctx.tenantId)) throw new Error("tenant not allowed");
if (Date.now() > Date.parse(ctx.expiresAt)) throw new Error("context expired");
if (!verifySignature(ctx)) throw new Error("context signature invalid");
return Object.freeze(ctx);
}

The tenant must come from a verified token, session, or server-side mapping. Ignore a tenant identifier supplied in user text, retrieved content, tool arguments, or model output. Those values can be useful hints for a clarification flow, but never for authorization.

Enforce authorization before retrieval

Microsoft’s secure multitenant RAG guidance recommends putting an API layer in front of the database or vector store. Every retrieval request should carry tenant and user authorization filters, and filtering must happen before grounding data reaches the model.

The unsafe sequence is:

search all tenants -> rank results -> filter unauthorized chunks -> prompt model

The safe sequence is:

verify identity -> construct trusted filters -> search only eligible chunks -> rank -> prompt model

Attach security metadata at ingestion time from controlled services, not from document text:

{
"chunk_id": "chunk_8841",
"document_id": "doc_220",
"tenant_id": "tenant_acme",
"allowed_subjects": ["group:finance", "user:usr_91"],
"classification": "confidential",
"source_system": "sharepoint",
"source_version": "v17",
"content_hash": "sha256:...",
"review_status": "approved"
}

Then make the filter mandatory in the retrieval API:

def retrieve(query: str, ctx: RequestContext, limit: int = 8):
filters = {
"tenant_id": ctx.tenant_id,
"allowed_subjects": {"overlap": [ctx.subject, *ctx.roles, *ctx.scopes]},
"review_status": "approved",
}

candidates = vector_store.search(
query=query,
filters=filters,
limit=limit,
)

if any(item.tenant_id != ctx.tenant_id for item in candidates):
raise SecurityError("retrieval boundary violation")
return candidates

Apply the same rule to hybrid search, reranking, parent-document expansion, citations, summaries, and cache reads. Filtering only the first vector query is not enough.

Isolate memory and caches

Memory is derived data, but it still carries tenant authority. A memory record should include tenant, subject scope, source, policy version, and lifecycle status.

memory_record:
id: mem_01JX
tenant_id: tenant_acme
subject_scope: [user:usr_91, workflow:billing]
kind: episodic
source_trace: tr_8841
policy_version: data-policy-v7
status: active
expires_at: 2026-11-10T00:00:00Z

Do not use a global cache key such as query_hash. Include authorization context and policy version:

cache_key = hash(
tenant_id,
subject_id,
role_set,
policy_version,
normalized_query,
retrieval_config_version
)

Invalidate memory and caches when a user leaves a group, a document is revoked, a tenant is suspended, or a policy changes. A tenant-safe system must protect deletion and revocation paths as carefully as ingestion.

Keep tool execution server-side

The model can propose an operation. The server must decide whether that operation is authorized for this caller and tenant.

AWS’s multitenant analytics reference uses three independent layers: cryptographically verified request identity, semantic validation before data access, and programmatic row-level isolation before the model receives data. The model sees only pre-filtered schemas or views, not raw tables.

async function callTool(name: string, args: unknown, ctx: RequestContext) {
const tool = toolRegistry.require(name);
const decision = await policy.check({
subject: ctx.subject,
tenant: ctx.tenantId,
roles: ctx.roles,
tool,
args,
});

if (decision.effect !== "allow") {
await audit.record({ ctx, tool: name, decision, argsHash: hash(args) });
throw new Error("tool denied");
}

return tool.invoke({ ...args, tenantId: ctx.tenantId, subject: ctx.subject });
}

Never let the model choose a different tenant by adding tenantId to tool arguments. The gateway overwrites scope with trusted context or rejects the call.

For shared MCP servers, propagate user identity and enforce it again at the backend. For sensitive or regulated systems, a tenant-specific MCP server or project boundary may reduce the identity-mapping and lateral-movement surface at the cost of more operations.

Test for cross-tenant leakage deliberately

A normal happy-path test proves very little. Build a tenant matrix with similar queries, similar documents, and deliberately colliding identifiers.

1. Insert semantically similar documents for tenants A and B.

2. Query tenant A using terms that strongly match tenant B’s document.

3. Attempt to override tenant scope in user text and tool arguments.

  1. Reuse a conversation, cache key, memory ID, and queue payload across tenants.

5. Revoke access after indexing and verify retrieval stops immediately.

  1. Run concurrent requests with different tenants and inspect every response and trace.
  2. Attempt parent-document expansion from an authorized chunk into an unauthorized document.
  3. Verify citations, summaries, rerankers, and fallback paths preserve the same filters.
assert all(result.tenant_id == ctx.tenant_id for result in results)
assert all(citation.tenant_id == ctx.tenant_id for citation in citations)
assert cache_key.includes(ctx.tenant_id)
assert tool_request.tenant_id == trusted_context.tenant_id

The important assertion is not only “the final answer contains no secret.” It is “unauthorized data was never a candidate for ranking, prompt construction, caching, logging, or citation.”

What practitioners are asking

An AI Agents community discussion describes the operational concern directly: shared worker pools, sensitive document extraction, iterative code loops, and fair-share policies can make per-tenant isolation difficult to scale. The answer is not necessarily one Kubernetes pod per tenant. It is a deliberate choice of isolation level at each layer, with server-side identity and policy enforcement that remains stable across shared infrastructure.

The secure RAG walkthrough linked below emphasizes that retrieval metadata should include tenant, owner, role, classification, provenance, and review status. It also highlights caches and summaries as derived-data leak paths. A tenant boundary that exists only in the primary database but not in memory, retrieval, or observability is incomplete.

References & Community Insights

Final checklist

Before launching a multi-tenant agent, verify:

  1. Is tenant identity derived from authenticated server-side context rather than prompt text?

2. Are retrieval filters applied before ranking and context assembly?

3. Do memory, summaries, citations, and caches carry tenant and subject scope?

4. Are tool calls authorized independently of model output?

5. Do shared MCP servers receive and enforce propagated user identity?

6. Are row-level or equivalent controls applied beneath the model?

  1. Can revocation, concurrency, and fallback paths be tested with a tenant matrix?

8. Do logs and traces avoid cross-tenant prompt and result leakage?

Multi-tenant agent security is not a single database setting. It is a chain of custody for identity and data. Preserve that chain from request admission to retrieval, memory, tool execution, and response—and the model can remain flexible without becoming the authority that decides whose data it may see.

Want to implement this in your business?

Mapki designs bespoke AI agents, custom workflow automations, and tool-agnostic integrations tailored specifically to your existing ERP, CRM, and databases.

Tags:Multi-Tenant AIAI Agent SecurityRAG SecurityAuthorizationAgent MemoryEnterprise AI