Flat Spell Technologies

Azure AI Foundry · FedRAMP Moderate engineering

Choose Azure AI Foundry models, deployment types, and pricing

Select an eligible model and processing boundary before optimizing price. A catalog listing, low token rate, or regional resource name does not establish where every request is processed.

1. Make an explicit model approval record

Record the provider offering, cloud, account kind, project API generation, model publisher, model name and version, deployment name, SKU, processing geography, feature status, and applicable package reference. A deployment name is a resource label, not proof of the underlying model version.

Distinguish models sold by Azure from other catalog offerings, marketplace providers, and externally hosted endpoints. Review each offering’s terms, data processing, and authorization scope. A model available in commercial Foundry is not necessarily available or acceptable in Azure Government.

Read the current Government model table for supported regions and deployment types. Do not select a model solely because it is described as open source or has no API-key requirement; hosting and operating it introduces separate runtime and supply-chain responsibilities.

2. Match deployment type to the processing requirement

Processing choices to verify against current model documentation
Deployment familyProcessing considerationReview question
Commercial GlobalProcessing may use regions beyond the resource’s own geographyIs that geographic processing scope allowed?
Commercial Data ZoneProcessing constrained to the documented data zoneDoes that zone match the approved data handling boundary?
Commercial Standard / regional provisionedCurrent documentation describes processing within the designated geography, including possible cross-region processing for operationsAre the geography and any operational movement acceptable?
Government Data ZoneDocumented processing within the USGov data zoneDoes the program permit that zone rather than one region?
Government Standard / regional provisionedGovernment deployment documentation describes regional processingIs this deployment available for the selected model and approved region?

These are routing considerations, not a list of authorized SKUs. Verify the selected service and deployment in the provider package. At-rest location and inference processing location are different facts; record both.

Microsoft’s privacy documentation also describes stateful data and abuse-monitoring behavior. Neither a private endpoint nor a client-side “do not store” option should be treated as a universal zero-retention guarantee. Fine-tuning, batch, Responses, and agent state add their own data flows and must be approved separately.

Azure Government has a distinct abuse-monitoring posture. The Government deployment guide says not all abuse-monitoring features are enabled, while automated content classification and filtering remain enabled by default. Customers are responsible for reasonable technical and operational measures to detect and mitigate prohibited use. Review those responsibilities rather than assuming the commercial-cloud monitoring behavior applies unchanged.

3. Estimate token charges with actual rates

Use the Azure OpenAI pricing information and your cloud account’s actual billing terms. Read rates for the actual model, deployment type, cloud, region, currency, and commercial agreement. Capture the date and billing unit. Cached-input rates, batch rates, image/audio units, and provisioned throughput are distinct mechanisms and should not be mixed into a single token rate.

from security_patterns import token_cost

estimated_model_charge = token_cost(
    input_tokens=measured_input_tokens,
    output_tokens=measured_output_tokens,
    input_per_million=approved_input_rate,
    output_per_million=approved_output_rate,
)

The calculator component uses decimal arithmetic and rejects invalid counts and rates. It covers uncached input/output token charges only. Supply current rates; no rate in this guide is an Azure price quote.

For RAG, measure system instructions, the user question, retrieved context, and generated output. For agents, count every model turn and tool-result context. A successful request that uses five turns can cost substantially more than a one-turn estimate. Bound retrieved context, output, retries, and iterations.

4. Include the supporting platform

  • Search capacity and ingestion, embeddings, and any reranking.
  • Agent state, file storage, and Cosmos DB throughput or request charges where applicable.
  • Application compute, approved container registries, and background workers.
  • Private endpoints, DNS/resolver infrastructure, firewall, approved connectivity, and data transfer.
  • Telemetry ingestion/retention, security tooling, backups, and recovery exercises.
  • Evaluation runs, release tests, engineering, and ongoing control maintenance.

Budget alerts notify; they are not a reliable hard stop on all usage. Implement application admission limits, per-user or per-tenant quotas, bounded concurrency, and retry ceilings. Reconcile measured usage with billed usage in your cloud account.

Compare provisioned capacity with measured peak demand, latency, model eligibility, and utilization. Reserved capacity can leave paid idle headroom. Evaluate cost alongside acceptance criteria rather than moving to an unapproved route to avoid throttling.

5. Govern upgrades, routing, and lifecycle

Test a model/version change against the same security and quality gates as an application release. Record the selected version and deployment configuration. Reevaluate grounding, extraction accuracy, tool behavior, content filtering, token usage, and latency before promoting it.

Model routers and fallback chains can widen the approved model set or processing boundary. Government documentation currently lists the model router as unavailable. Even where supported, every routing target must be approved; avoid uncontrolled fallback to public or third-party inference.

Plan for model retirement, quota changes, service incidents, and insufficient capacity. A safe fallback can be a queue, a limited non-AI path, or an explicit service-unavailable response—not an unapproved model. Continue with release evaluation and the relevant architecture guide.

Sources and technical review

Technical review: . Microsoft documentation changes over time; recheck feature availability and your authorization package before deployment.

Confirm the offering and applicable authorization status in the FedRAMP Marketplace and review the provider package and your system security plan. Service availability is not an authorization determination.