Azure Data Factory · FedRAMP Moderate engineering
Monitor Azure Data Factory pipelines and control performance and cost
Monitor pipeline success, data completeness, runtime health, and cost as separate signals. A green pipeline status is useful, but it does not prove the intended data arrived safely or on time.
1. Define operational and data-quality acceptance criteria
Define freshness lag, completion deadline, reconciled row/key counts, schema contract, recovery time, and acceptable replay behavior per ingestion source. Track a sanitized batch identifier, pipeline run ID, committed cursor, and approved destination. Keep the diagnostic workspace, alert receivers, and evidence retention in the system boundary with appropriate access and deletion rules.
Distinguish orchestration failure from data-quality failure and missed execution. An empty copy can succeed even when the source stopped producing records; a successful activity can publish the wrong schema. Alerts need both platform failures and a source-specific expected arrival or checkpoint-age signal.
Run tests with synthetic data first. Audit payload minimization in pipeline parameters, activity input/output, SQL text, exception messages, file names, and external compute logs. The provider offering’s public scope does not remove your responsibilities for log handling and access.
2. Configure resource-specific diagnostic tables
In the factory’s Diagnostic settings, choose the log categories your workload needs, such as PipelineRuns, ActivityRuns, and TriggerRuns, and send them to an approved Log Analytics workspace. Microsoft recommends resource-specific mode, which produces ADFPipelineRun, ADFActivityRun, and ADFTriggerRun tables instead of a shared AzureDiagnostics table. Include relevant metrics; do not enable every SSIS category for a factory without SSIS workloads.
Discover the categories actually available on the resource before configuring them:
az monitor diagnostic-settings categories list --resource "$FACTORY_RESOURCE_ID"
Protect workspace reader access, alert detail, exports, and retention. Microsoft notes that diagnostic ingestion may take up to 15 minutes, so account for collection delay when choosing alert windows. Check that all expected tables receive synthetic test events; the existence of a diagnostic setting is not evidence that events are arriving.
This resource-specific KQL example summarizes terminal pipeline failures without projecting activity payloads or error messages. Confirm column names and field availability in the target workspace before saving the rule:
ADFPipelineRun
| where TimeGenerated >= ago(30m)
| where Status == "Failed"
| summarize FailedRuns = dcount(RunId) by _ResourceId, PipelineName
| where FailedRuns > 0
Use a separate heartbeat/checkpoint-age check for runs that never start or never finish, and a data-quality alert for failed publication checks. Deduplicate alerts by source/batch or run ID and avoid copying sensitive error strings into email or incident channels.
3. Apply secure activity policies and keep minimal telemetry
For activities whose inputs or outputs can contain sensitive values, configure secure input/output policy and test the monitoring views and diagnostic exports. These flags hide the corresponding activity input/output in monitoring; they do not guarantee that parameters, errors, connector traces, or every downstream log is redacted.
{
"policy": {
"timeout": "00:30:00",
"retry": 2,
"retryIntervalInSeconds": 60,
"secureInput": true,
"secureOutput": true
}
}
Keep a separate sanitized operational record with batch ID, schema version, start/end times, counts, validation outcome, and committed cursor. Hiding the Copy output can hide useful counters too, so obtain required reconciliation metrics through the approved validation/control-store path. Do not re-expose the entire output merely to restore a dashboard.
Choose timeout and bounded retry for each failure mode. Credential denial, invalid schema, and denied endpoints usually need correction rather than repeated load. Retries should reuse an immutable batch contract and deterministic target operations.
4. Tune the actual bottleneck and runtime model
Measure source query duration, runtime queue time, transfer throughput, sink throttling, startup delay, and validation time separately. For an Azure IR Copy activity, data integration units (DIUs) describe allocated transfer capacity. Self-hosted IR performance depends on its own hosts and concurrent-job settings; increasing an Azure DIU setting is not a way to resize those nodes.
Increase parallel copies, partitioning, or ForEach concurrency only after validating source and destination limits and the replay contract. Concurrent extraction can overload a database or make a weak timestamp contract less reliable. Apply source indexes and approved partition predicates before simply multiplying workers.
Managed-VNet IR can incur cold-start delay. Microsoft documents Copy-compute TTL and warns that reserved compute/TTL affects billing; the per-activity output may not contain a billingReference in TTL scenarios. Interactive authoring TTL and Copy-compute TTL are distinct. Mapping Data Flows use managed Spark compute and a different vCore meter, including debugging time.
Benchmark with synthetic or approved nonproduction data, retain the same volume and schema between comparisons, and measure cost per validated published batch rather than throughput alone. Recheck regional/feature availability in Azure Government before selecting an optimization.
5. Estimate cost using the right units
Use current rates from the selected cloud, region, currency, agreement, and pricing calculator. This guide quotes no Azure prices. The public pricing page distinguishes orchestration/activity runs, runtime-based execution, data movement, Data Flow execution/debugging, and factory operations. Integration-runtime execution is prorated by the minute and rounded up; apply the actual meter’s rules rather than one rounding rule to every charge.
| Component | Unit basis | What to include |
|---|---|---|
| Orchestration | Activity runs / 1,000 × applicable rate | Retries and orchestration activities, not only pipeline starts |
| Azure Copy data movement | DIUs × billed hours × rate per DIU-hour | Meter rounding and managed-VNet/TTL reservations |
| Self-hosted execution | Applicable execution meter + host costs | VMs or on-premises hosts, disks, management, and redundancy |
| Mapping Data Flow | vCores × billed hours × rate | Execution, debug sessions, and applicable reservations/TTL |
| Supporting services | Service-specific meters | Storage operations/capacity, SQL, logs, private networking, Key Vault, and data transfer |
For unit arithmetic only, a hypothetical non-TTL Copy using 8 DIUs for 2 minutes 20 seconds would use 3 billed minutes under the documented minute-rounding example: 8 × 3 / 60 = 0.4 DIU-hours. Multiply by the applicable data-movement rate, then add orchestration and supporting services. This is not an Azure price quote or a universal reservation-billing formula.
Allocate cost by environment and ingestion source, set budgets, and reconcile estimates with actual consumption and billing after the pilot. Include idle debug sessions, repeated validation attempts, retained quarantined batches, and cloud egress where applicable.
6. Exercise alerts, redaction, recovery, and budgets
- Inject a controlled connector failure and verify a payload-free alert reaches the approved route.
- Suppress a scheduled input and verify the freshness/heartbeat check catches the missing data.
- Reject a malformed schema even when Copy succeeds, keeping the cursor uncommitted.
- Search the monitoring UI, exported logs, and alert content for a synthetic sensitive marker; confirm the intended redaction boundary.
- Stop one self-hosted node and verify capacity and recovery alerts while testing safe replay.
- Compare measured meter units and total supporting-service costs with the pilot estimate.
Save diagnostic configuration, role scopes, log/alert samples, acceptance timings, and cost assumptions as assessment evidence. Revisit checkpoint and publication recovery and release gates when a change alters these operational expectations.
Sources and technical review
Technical review: . Recheck cloud, regional, connector, and authorization scope before implementation.
- Resource-specific diagnostic tables and ingestion delay
- Activity policies and concurrency
- Copy activity performance and tuning
- Managed-VNet runtime compute and TTL billing
- Mapping Data Flow performance and compute
- Data Factory pricing and meter units
- ADFPipelineRun table schema
- Pipeline ARM schema and secure activity policy
Public offering and service-scope checks are documented in the Data Factory hub. The protected provider authorization package and live Azure behavior have not been reviewed or tested by these guides.