Interactive Benchmark Skill
Know your query will perform before it hits production
Customers want to prove
Interactive works for them
Interactive Warehouses are designed for high-concurrency, low-latency workloads — dashboards, APIs, real-time apps.
Customers are excited about the promise. But before committing, they need to answer one question:
Run their actual queries under realistic concurrent load to validate latency targets.
P95 latency under N concurrent users — measured, not estimated.
Which warehouse size and cluster count will deliver the SLA without overspending?
Performance testing
is painful
Latency guarantees
Will this query run under 200ms with 50 concurrent users? Customers can't answer this without a real test.
Sizing guesswork
What warehouse size? How many clusters? Trial-and-error is expensive and slow.
Complex setup
Standing up a proper load-test environment with realistic concurrency is tricky to get right.
What customers face today
tuning
No clear path to optimize a query for interactive latency targets. Manual EXPLAIN plans, guessing at clustering keys, iterating blindly.
testing
Setting up a load test is the hardest part. Deploying Locust, configuring compute pools, warming caches, collecting server-side metrics.
sizing
Over-provision or under-deliver. Without iterative testing, customers either waste credits on oversized warehouses or miss their SLA.
Just say what you need
Manually deploy Locust, configure SPCS compute pools, write test scripts, warm caches, collect metrics, resize warehouses, repeat.
"I want this query to run under 200ms with 50 concurrent users." The skill handles everything else.
Automated end-to-end
Suitability gate
Validates the query is a good fit for interactive workloads before spending any credits on load testing.
Full load test
Deploys Locust + FastAPI on SPCS, warms caches, runs baseline and real tests with your concurrency target.
Iterative scaling
If P95 misses the goal, it automatically scales out clusters or up warehouse size within your limits.
Guardrails
Set max warehouse size and cluster count so costs stay controlled
Query variations
Tests a single query with variations or a set of provided queries
HTML report
P50 / P95 / P99 latencies, throughput, and optimization recommendations
Three phases, fully automated
SPCS deployment
Locust
1 instance
autostart mode
Phase 1 Baseline
Phase 2 Load test
N concurrent users
LOAD BALANCER
Benchmark API #1
FastAPI · 4 workers · 50 conn
Benchmark API #2
FastAPI · 4 workers · 50 conn
Benchmark API #3
FastAPI · 4 workers · 50 conn
OAuth token
QUERY_TAG · no cache
Interactive
Warehouse
Multi-cluster
1 – N clusters
USE_CACHED_RESULT = False
Standard scaling policy
Interactive tables or zero-copy
Results Collection & Escalation Loop
Client-side (Locust)
HTTP timing · P50 / P95 / P99 · Failure rate · RPS
Server-side (Snowflake)
QUERY_HISTORY · Execution time · Queue wait · Cluster usage
Goal Check
P95 ≤ target? → Report.
Missed? Scale & re-test.
You stay in control
- Set a latency target — "Run under 200ms at P95" and the skill works toward that goal
- Cap warehouse size — won't go beyond the max size you approve (e.g. no 3XL)
- Cap cluster count — controls multi-cluster scaling to keep costs predictable
- Automatic cache warming — resets and warms cache between each iteration so results are reliable
- Clean up when done — offers full teardown of SPCS services and compute pools
What the skill handles for you
Across 46 authored files
Across 3 phases
snowflake-interactive, html-authoring
CoCo tools, bash, docker, SQL
Gates, conditionals, 3-way splits
User confirmations + hard-stop gates
Branches minus merge points
Longest path with 5 escalations
2 services + 2 compute pools
User-configurable (default 5)
Sizing, validation, troubleshooting
Templates, configs, scripts, reports
Demo
Lessons learned
Natural language
for flexibility
Define the workflow in natural language. It gives the AI agent the flexibility to adapt, recover from errors, and handle edge cases that rigid code cannot anticipate.
Deterministic tools
for predictability
Wrap key operations in scripts and tools. The more deterministic tools you give the agent, the easier it is to define the workflow and get predictable, repeatable outcomes.
Tools are as important
as the skill itself
A skill is only as good as the tools it can call. Invest in building reliable, well-scoped tools — they are the foundation the entire skill definition rests on.
Thank you
Interactive Benchmark Skill