Connect Your Production, Simulate Against It: The Auto-Calibration Engine
Import real metrics from Datadog, real incidents from Jira, real infrastructure from Terraform — and the simulation automatically matches your production behavior.
The Gap Between Simulation and Reality
Most architecture simulations use generic numbers. "Assume 100ms latency." "Assume 99.9% availability." "Assume 1000 req/s."
But your production doesn't work in round numbers. Your real P99 is 347ms. Your real peak is at 2:15 PM on Tuesdays. Your payment service fails 0.3% of the time, but only on international transactions.
What if your simulation used your real production data?
The Auto-Calibration Pipeline
PraxiRun connects to your existing monitoring, incident management, and infrastructure tools — then automatically calibrates the simulation to match your production behavior.
Step 1: Connect Your Monitoring
Connect your Datadog, New Relic, CloudWatch, or Prometheus instance. PraxiRun imports the last 30 days of:
- Traffic patterns — when are your peaks? What's the diurnal curve?
- Latency distributions — P50, P95, P99 per service, not averages
- Error rates — which services fail? How often? With what patterns?
- Resource utilization — CPU, memory, connection pools per service
Step 2: Connect Your Incident History
Connect Jira or ServiceNow. PraxiRun imports your incident records and extracts:
- Failure frequency — "payment service had 3 incidents last month"
- Mean Time To Recovery — "average MTTR is 45 minutes"
- Root causes — "database connection exhaustion caused 60% of incidents"
- Impact patterns — "incidents during peak hours affect 10x more users"
These become realistic failure scenarios in your simulation. Not generic "service down" scenarios — YOUR actual failure modes.
Step 3: Connect Your Infrastructure
Connect your Terraform state or Kubernetes API. PraxiRun imports:
- Actual resource configuration — instance types, replica counts, HPA settings
- Network topology — VPCs, subnets, security groups
- Scaling rules — min/max replicas, CPU thresholds, cooldown periods
Your simulation now has the exact same infrastructure as production.
Step 4: Auto-Calibration
The calibration engine combines all three sources and produces a SimCalibrationProfile:
Traffic: diurnal pattern, base=45 rps, peak=380 rps (Tuesdays 2:15 PM)
Latency: P50=23ms, P95=89ms, P99=347ms (from Datadog)
Error rate: 0.3% (payment service, international only)
Replicas: 3 (from Terraform state)
Scenarios: "Payment timeout" (3/month, MTTR 45min)
"DB connection exhaustion" (1/month, MTTR 90min)
Cost: $12,400/mo (from billing)
Availability: 99.92% (calculated from real error rate)
Your simulation now behaves like your production. Every latency, every error rate, every scaling behavior — calibrated from real data.
What You Can Do With a Calibrated Simulation
"What if we add a cache?"
Before: you guessed it would help. Maybe 30% improvement?
After: simulation shows P99 drops from 347ms to 52ms. DB utilization drops from 78% to 22%. Cache hit rate stabilizes at 73% after warmup. Cost increases by $45/mo.
You have the exact numbers. Not estimates — data.
"What if we lose an availability zone?"
The simulation kills all services in AZ-1a (matching your real Terraform topology). Traffic redistributes to AZ-1b and AZ-1c. Simulation shows:
- 23 seconds of elevated error rate during failover
- P99 spikes to 890ms for 45 seconds
- Auto-scaling adds 2 replicas in 90 seconds
- Full recovery in 2 minutes
You know exactly what happens. Before it happens.
"What if Black Friday traffic hits?"
Multiply your real peak traffic by 5x. Simulation shows:
- Database connection pool exhausts at 3.2x current peak
- Payment service becomes the bottleneck at 4.1x
- Cache hit rate drops to 45% (cold cache for new product pages)
- Breaking point: 4.8x current peak
You know your limits. And you know exactly which component to scale.
30 Connectors
PraxiRun connects to your entire stack:
| Category | Connectors | |---|---| | Monitoring | Datadog, New Relic, Dynatrace, CloudWatch, Prometheus, Elasticsearch, Grafana | | Incidents | Jira, ServiceNow, PagerDuty, OpsGenie | | Infrastructure | AWS, Azure, GCP, Kubernetes, Terraform, Istio | | Physical | MQTT, ROS 2, OPC-UA, MAVLink, HL7 FHIR, BACnet | | CI/CD | GitHub, Jenkins | | Communication | Slack, Email (Resend) |
Connect once. Calibrate continuously. Simulate with confidence.
Get Started
- Open PraxiRun
- Go to Connectors → connect your Datadog/CloudWatch
- Click "Import Production Data" in the dashboard
- The simulation auto-calibrates
- Run what-if scenarios against your real baseline
Your architecture decisions are now backed by data, not debate.
Ready to test your architecture skills?
Try a Free Simulation →