Flat Spell Technologies

Azure AI Foundry · FedRAMP Moderate engineering

Evaluate and monitor Azure AI Foundry configurations

AI readiness includes access control, data handling, failure behavior, and operational evidence as well as answer quality. Use deterministic security tests alongside task-specific evaluation.

1. Build a controlled evaluation dataset

Use synthetic or explicitly approved test data. Separate RAG questions with expected source documents, extraction fixtures with known fields and pages, and agent requests with permitted and forbidden actions. Include multiple tenants and users with different permissions.

Keep dataset versions, expected outcomes, model deployment/version, prompt revision, retrieval configuration, and evaluator version in the result manifest. An evaluation using a model as judge introduces another data processor; approve that model and its endpoints as well.

Microsoft’s current Government platform table lists built-in evaluations and optimization as unavailable, while prompt-agent tracing is available. Use an approved application-side evaluation harness where the built-in feature is unsupported. Do not upload regulated evaluation data to a commercial workspace as a workaround.

2. Define release gates that can fail

Example security and task acceptance checks
ScenarioExpected behaviorEvidence
Cross-tenant retrievalNo unauthorized chunk reaches the modelFilter, context manifest, denied result
Prompt injectionRetrieved text cannot widen permissions or authorize toolsTool/authorization decisions under malicious fixtures
Forbidden tool callDispatcher denies before backend executionDenied event and zero backend calls
Extraction mismatchInvalid amount, page, or source ID fails reviewValidator output and routed exception
Public network callerDenied even with a valid scoped identityNetwork decision and resource configuration
Revoked accessNext request observes the defined revocation policyLifecycle timestamps and denied action

Define task-quality thresholds against your own acceptable error rate. RAG needs citation support and authorized source coverage; extraction needs field accuracy and review rates; agents need permitted-action success and denied-action correctness. A model-generated score alone should not overrule deterministic access checks.

3. Log decisions and metrics without default content capture

Collect latency, token counts, model/version, deployment identifier, retrieval count, tool name, allow/deny decision, error class, and an approved correlation identifier. Treat user and document identifiers according to the program’s privacy policy.

{
  "event": "ai_request_completed",
  "correlation_id": "synthetic-run-001",
  "deployment": "approved-deployment-name",
  "retrieved_chunks": 3,
  "tool_decision": "denied",
  "input_tokens": 1400,
  "output_tokens": 180,
  "latency_ms": 920
}

This is an example event, not measured production performance. Omit prompts, retrieved text, document contents, raw model outputs, tokens, and secrets by default. Inspect SDK debug logging, OpenTelemetry exporters, evaluation integrations, exception handlers, and tool tracing: content may be captured outside your application’s main logger.

In Azure Government, Microsoft documents that not all provider abuse-monitoring features are enabled. Include the customer’s required detection and mitigation measures in the operating model; automated content classification and filtering remain enabled by default. Review the Government data-handling distinction when choosing telemetry and response procedures.

If approved content tracing is necessary, restrict it to the reviewed scope and retention, protect access, and record the purpose. Document provider-side state and abuse monitoring separately; telemetry redaction does not control those stores.

4. Verify alerts, limits, and recovery behavior

Monitor identity failures, denied data/tool actions, unusual request volume, repeated quota failures, latency, retry amplification, and failures in ingestion or deletion. Use context to avoid paging on every expected denial, but retain reviewable evidence.

Test the complete alert path with synthetic events, including routing and escalation. Exercise a model outage, unavailable Search service, a dead tool backend, expired credentials, and missing DNS. The application should fail within bounded time rather than retry forever or fall back to an unapproved endpoint.

Rehearse backup/restore for required stores, confirm access controls after recovery, and reapply retained deletion instructions according to your policy. A backup existing does not prove restore readiness.

5. Package evidence with configuration provenance

Retain approved infrastructure definitions, account and network configuration exports, identity assignments, tool allowlists, data-store lifecycle policies, model records, dataset hashes or version identifiers, and acceptance results. Link the artifacts to the actual environment and release.

Keep failed, passed, skipped, and unrun checks distinct. A scan that could not authenticate is not a passing configuration check. Sign or integrity-protect evidence as required, restrict modification, and apply the approved retention policy.

These artifacts can support AC, IA, SC, AU, CM, CA, SI, and CP control implementation review. Mapping a test to a control does not establish complete coverage or replace assessment and authorization decisions.

6. Repeat targeted checks when the system changes

Rerun affected evaluations after changing a model/version, prompt, retrieval index, ACL mapping, tool schema, SDK, cloud feature, or routing behavior. Preserve the prior results for comparison. Use a limited rollout only where the program permits it and acceptance criteria are explicit.

The application guardrails have offline tests for empty permissions, cross-tenant retrieval, invalid IDs, forbidden tools, extraction validation, and token-cost units. Those tests do not exercise Azure, real identity tokens, private endpoints, or provider retention.

Return to the configuration hub for the full readiness sequence or use RAG-specific checks and agent tool checks to define your test fixtures.

Sources and technical review

Technical review: . Microsoft documentation changes over time; recheck feature availability and your authorization package before deployment.

Confirm the offering and applicable authorization status in the FedRAMP Marketplace and review the provider package and your system security plan. Service availability is not an authorization determination.