Azure AI Foundry · FedRAMP Moderate engineering
Build a permission-aware Azure AI Foundry RAG architecture
Retrieval-augmented generation should retrieve only documents the authenticated user is allowed to read. Enforce that boundary before a chunk reaches the model, rather than asking the model to keep secrets.
1. Separate ingestion, retrieval, and inference
- Ingestion identity
Approved source → extraction → chunks + source ACLs - Private Search index
Tenant, group permissions, source ID, and content on every chunk - Application retrieval
Verified user → server filter → authorization recheck - Approved model
Minimal authorized context → answer + checked source citations
This guide uses application-orchestrated RAG. It avoids assuming that a managed retrieval connector provides the same per-user authorization as the source. Approve the inference model, embedding model if used, Search service, storage, extraction service, and application runtime independently.
Provision the private network foundation first. Use a write-capable ingestion identity and a retrieval identity with only the permissions it needs. Model inference should not grant the user direct Search administrator or storage credentials.
2. Carry access metadata onto every chunk
At ingestion, resolve source permissions through a trusted connector. Write tenant and allowed-group identifiers alongside every derived chunk. Keep a source version, chunk ID, and the source-to-chunk relationship in your ingestion records so updates and deletions can be propagated.
{
"name": "approved-document-chunks",
"fields": [
{"name":"id","type":"Edm.String","key":true},
{"name":"tenant_id","type":"Edm.String","filterable":true},
{"name":"allowed_group_ids","type":"Collection(Edm.String)","filterable":true},
{"name":"content","type":"Edm.String","searchable":true}
]
}
View the full lexical Search index example on GitHub. Its ACL fields are retrievable for a server-side authorization recheck; never send them directly to the browser. If your design keeps ACL fields nonretrievable, recheck against a separate authoritative permission store instead.
This minimal index is full-text retrieval, not a vector-index configuration. To add vector search, choose an approved embedding deployment, set dimensions to that model’s actual output, configure a supported vector algorithm/profile, and preserve the same authorization filter. Embeddings and intermediate chunks remain sensitive derived data.
Add vector or hybrid retrieval with the same permission boundary
+Use an approved embedding deployment for both ingestion and queries. Record the exact model/version and dimension count. The model's commercial availability does not establish Government availability or authorization scope.
View the vector-index generator on GitHub and the lexical index file together. Run python3 build_vector_index.py --dimensions YOUR_APPROVED_DIMENSION_COUNT to generate an index with a cosine HNSW profile and preserved ACL fields. Replace the dimension placeholder with the actual embedding output size, verify service/API support, and deploy the resulting index through your approved index-management identity. The generator does not create a resource or call Azure.
Store one approved embedding per chunk in content_vector. Reject vectors with the wrong size; update embeddings and the index together when changing the model or dimensions. The schema keeps vectors nonretrievable.
from azure.search.documents.models import VectorizedQuery
# query_embedding comes from the same APPROVED deployment/dimensions as ingestion.
vector_query = VectorizedQuery(
vector=query_embedding,
k_nearest_neighbors=50,
fields="content_vector",
)
results = search_client.search(
search_text=question,
vector_queries=[vector_query],
vector_filter_mode="preFilter",
filter=filter_expression,
top=5,
select=["id", "content", "tenant_id", "allowed_group_ids"],
)
# Recheck result permissions before constructing model context.
This hybrid query combines the text query and vector query. It assumes an initialized Search client and a trusted filter expression from the next step. Microsoft recommends preFilter for selective authorization filters because it avoids the recall problems that can occur when filtering only after candidate selection. Confirm index/API compatibility; do not retry a failed filtered query by removing the filter.
The result authorization recheck remains required in this example. Test permissions for full-text, vector, hybrid, and any semantic reranking path. A relevance feature must not create an alternate route around document access controls.
3. Build the filter from verified user identity
Your authentication middleware must validate token signature, issuer, audience, tenant, expiry, and required application roles. Resolve group overage through an approved identity path. Do not accept tenant or group IDs from an untrusted request body, and do not drop the filter when group resolution fails.
from security_patterns import acl_filter, select_authorized_chunks
# trusted_identity is produced by authenticated server middleware.
filter_expression = acl_filter(
trusted_identity.tenant_id,
trusted_identity.verified_group_ids,
)
results = search_client.search(
search_text=question,
filter=filter_expression,
top=5,
select=["id", "content", "tenant_id", "allowed_group_ids"],
)
chunks = select_authorized_chunks(
results,
trusted_identity.tenant_id,
trusted_identity.verified_group_ids,
)
if not chunks:
return {"answer": "No authorized source material was found."}
The snippet belongs inside an authenticated application request handler; it is not a complete server. The guardrail module on GitHub validates UUID identifiers, includes a tenant predicate, denies empty permissions, and rechecks results before context creation.
Microsoft’s security-filter pattern compares strings; the Search service does not validate that your filter represents the real user. The application owns that decision. A managed identity’s ability to query an index is not permission for every end user to see every indexed document.
4. Construct minimal context and verify citations
Build a context envelope from authorized chunk IDs and text, capped by an application token budget. Delimit retrieved content as untrusted source material. Instruct the model to answer from supplied sources and return source IDs, but enforce the final citation allowlist in code.
A response that cites a source ID not present in the authorized context should fail validation or be regenerated under your approved policy. A syntactically correct citation still needs a groundedness check: the cited source must support the claim.
Do not use a system prompt as a document-access control. Treat instructions embedded in documents as source text, never as permission to add tools, visit URLs, or disclose other records. Keep tool selection and retrieval authorization outside the model.
5. Manage changes, deletion, caching, and diagnostics
Permission changes must invalidate or update affected chunks before they are used again. Define the latency your program accepts for ACL updates and test it. Document deletion must remove indexed chunks, embedding artifacts, caches, and any stateful conversation content according to the retention policy.
Cache keys must include the tenant and authorization context, not only the question text. A shared answer cache can leak another user’s sources even when the Search query is filtered correctly. Evaluate any semantic cache against the same rule.
Log correlation IDs, document IDs as permitted, access decisions, retrieved chunk count, latency, and model deployment identifiers. Avoid default prompt, chunk, and response capture in production telemetry. If content capture is necessary for review, put it behind explicit approval, access controls, and retention.
6. Prove the authorization boundary
- User A cannot retrieve user B’s unshared document, including by exact title or distinctive phrase.
- A matching group in a different tenant cannot retrieve the document.
- Empty, malformed, and unresolved group lists deny access rather than removing the filter.
- Revoked source permissions stop retrieval within the documented update interval.
- Injected document instructions cannot cause an external tool call or broaden the retrieval scope.
- Generated citations are a subset of authorized chunks and support the answer.
- Document deletion removes the relevant derived copies and caches under the approved lifecycle.
Record index configuration, source ACL mapping, identities and role scopes, denied-access tests, and lifecycle tests as supporting evidence. Continue with evaluation and monitoring. For managed agent retrieval, also review agent stores and tool authorization.
Sources and technical review
Technical review: . Microsoft documentation changes over time; recheck feature availability and your authorization package before deployment.
- Azure AI Search vector filtering
- Create a vector search index
- Azure AI Search document security filters
- Data, privacy, and security for models sold by Azure
- Configure Foundry Private Link
- Models sold by Azure in Azure Government
Confirm the offering and applicable authorization status in the FedRAMP Marketplace and review the provider package and your system security plan. Service availability is not an authorization determination.