Interactive Benchmark Skill

Know your query will perform before it hits production

Davide Mauri
Staff Product Manager
The Opportunity

Customers want to prove
Interactive works for them

Interactive Warehouses are designed for high-concurrency, low-latency workloads — dashboards, APIs, real-time apps.

Customers are excited about the promise. But before committing, they need to answer one question:

Can it sustain my workload at my scale?
They want to stress test it

Run their actual queries under realistic concurrent load to validate latency targets.

They want proof, not promises

P95 latency under N concurrent users — measured, not estimated.

They want the right sizing

Which warehouse size and cluster count will deliver the SLA without overspending?

The Challenge

Performance testing
is painful

Latency guarantees

Will this query run under 200ms with 50 concurrent users? Customers can't answer this without a real test.

Sizing guesswork

What warehouse size? How many clusters? Trial-and-error is expensive and slow.

Complex setup

Standing up a proper load-test environment with realistic concurrency is tricky to get right.

Why It's Hard

What customers face today

Query
tuning

No clear path to optimize a query for interactive latency targets. Manual EXPLAIN plans, guessing at clustering keys, iterating blindly.

Load
testing

Setting up a load test is the hardest part. Deploying Locust, configuring compute pools, warming caches, collecting server-side metrics.

Right-
sizing

Over-provision or under-deliver. Without iterative testing, customers either waste credits on oversized warehouses or miss their SLA.

The Solution

Just say what you need

Before

Manually deploy Locust, configure SPCS compute pools, write test scripts, warm caches, collect metrics, resize warehouses, repeat.

After

"I want this query to run under 200ms with 50 concurrent users." The skill handles everything else.

You set the goal. The skill figures out how to get there.
How It Works

Automated end-to-end

Suitability gate

Validates the query is a good fit for interactive workloads before spending any credits on load testing.

Full load test

Deploys Locust + FastAPI on SPCS, warms caches, runs baseline and real tests with your concurrency target.

Iterative scaling

If P95 misses the goal, it automatically scales out clusters or up warehouse size within your limits.

Guardrails

Set max warehouse size and cluster count so costs stay controlled

Query variations

Tests a single query with variations or a set of provided queries

HTML report

P50 / P95 / P99 latencies, throughput, and optimization recommendations

The Workflow

Three phases, fully automated

Phase 1 — Gather Inputs
Collect parameters — database, warehouse, latency goal, concurrency, max sizes
Present summary — all values in a single confirmation table
User confirms — nothing runs until you approve
Gate: All inputs confirmed
Phase 2 — Validate Suitability
Analyze the query — determine zero-copy vs. interactive tables approach
Run on both warehouses — standard vs. interactive, measure speedup
Suitability decision — stops early if query >10s or no speedup
Gate: <5s interactive, measurable speedup
Phase 3 — Run Benchmark
Deploy to SPCS — FastAPI + Locust on dedicated compute pools
Warm cache & baseline — reliable results, no cold-cache noise
Load test — real concurrency, P50/P95/P99 latencies
Escalate if needed — auto scale-out/up, re-test within limits
HTML report — results, recommendations, cleanup options
Goal: P95 meets your latency target
Architecture

SPCS deployment

Locust Pool

Locust

1 instance

autostart mode

Phase 1 Baseline
Phase 2 Load test
N concurrent users

LOAD BALANCER

API Pool · CPU_X64_M

Benchmark API #1

FastAPI · 4 workers · 50 conn

Benchmark API #2

FastAPI · 4 workers · 50 conn

Benchmark API #3

FastAPI · 4 workers · 50 conn

OAuth token
QUERY_TAG · no cache

Interactive
Warehouse

Multi-cluster

1 – N clusters

USE_CACHED_RESULT = False
Standard scaling policy
Interactive tables or zero-copy

Results Collection & Escalation Loop

Client-side (Locust)

HTTP timing · P50 / P95 / P99 · Failure rate · RPS

Server-side (Snowflake)

QUERY_HISTORY · Execution time · Queue wait · Cluster usage

Goal Check

P95 ≤ target? → Report.
Missed? Scale & re-test.

DNS
SQL
metrics
results
Guardrails

You stay in control

  • Set a latency target — "Run under 200ms at P95" and the skill works toward that goal
  • Cap warehouse size — won't go beyond the max size you approve (e.g. no 3XL)
  • Cap cluster count — controls multi-cluster scaling to keep costs predictable
  • Automatic cache warming — resets and warms cache between each iteration so results are reliable
  • Clean up when done — offers full teardown of SPCS services and compute pools
Under the Hood

What the skill handles for you

~4,000
Total Lines

Across 46 authored files

16
Workflow Steps

Across 3 phases

2
Sub-Skills

snowflake-interactive, html-authoring

26
Tools & Scripts

CoCo tools, bash, docker, SQL

36
Decision Branches

Gates, conditionals, 3-way splits

9
Stopping Points

User confirmations + hard-stop gates

~32
Cyclomatic Complexity

Branches minus merge points

48
Max Graph Depth

Longest path with 5 escalations

4
SPCS Resources

2 services + 2 compute pools

N
Escalation Loops

User-configurable (default 5)

11
Reference Docs

Sizing, validation, troubleshooting

27
Files Managed

Templates, configs, scripts, reports

Live

Demo

Takeaways

Lessons learned

01

Natural language
for flexibility

Define the workflow in natural language. It gives the AI agent the flexibility to adapt, recover from errors, and handle edge cases that rigid code cannot anticipate.

02

Deterministic tools
for predictability

Wrap key operations in scripts and tools. The more deterministic tools you give the agent, the easier it is to define the workflow and get predictable, repeatable outcomes.

03

Tools are as important
as the skill itself

A skill is only as good as the tools it can call. Invest in building reliable, well-scoped tools — they are the foundation the entire skill definition rests on.

Thank you

Interactive Benchmark Skill