Observability Treasure Hunt Print 24 copies, one per team. The facilitator drives the screen and teams answer from the projector. ← Hub

Observability Treasure Hunt

Team: ______________________    Total points: ______ / 150

Rules: Answer as many as you can in 35 minutes. Every answer needs proof: a screenshot, a number, or a service name. Bonus questions are worth double. Sharing clues with other teams costs 20 points.

Part A — Dynatrace

Use the environment shown by the facilitator. Service names differ between environments; any sensible answer with proof counts.

1
Top offender: Open Services. Which service has the highest failure rate in the last 2 hours? What's the rate?
10
2
Slowest service: Which service has the worst response time? Is it median or p95/p99 that looks bad?
10
3
Follow the ticket: Open Distributed Traces and pick a failed request. How many spans does it have? Which span took longest?
15
4
Who calls whom? From a service, look at its incoming and outgoing calls (service flow / Smartscape). Name 2 services it depends on.
10
5
Problem solver: Open Problems. Pick one problem. What's the root cause entity, and how many entities were affected?
15
6
Error detective: In Logs or a Notebook, run this. When was the error peak?
fetch logs
| filter loglevel == "ERROR"
| makeTimeseries errors = count(), interval: 5m
10
7
Saturation check: Is any host above 80% CPU?
timeseries cpu = avg(dt.host.cpu.usage), by: { dt.entity.host }
10
8
Bonus: Write your own DQL query that lists the top 5 slowest services by p95. Hint: fetch spans, summarize, percentile(duration, 95), sort, limit.
20

Part B — Azure (free tier)

The facilitator performs the steps in the Azure portal; teams answer from the screen and generate traffic from their phones. Everything is deleted at the end.

9
Build it: Create resource group rg-chaicart-demo in Central India, then a Web App. What's the public URL?
5
10
Crowd test: Everyone opens the URL on their phone and refreshes 10 times. In Application Insights → Live Metrics, what's the peak requests/sec?
10
11
Metric explorer: In Monitor → Metrics, chart “Requests” and “Response time” for the web app. What's the average response time?
5
12
Break it: The facilitator requests a URL that doesn't exist (e.g. /chai-not-found) 20 times. Where do the failed requests show up in Application Insights? Which status code?
10
13
KQL taste: In Application Insights → Logs, what does this return?
requests
| summarize count() by resultCode, bin(timestamp, 1m)
| order by timestamp desc
10
14
Money guard: In Cost Management → Budgets, create a monthly budget with an alert at 80%. Why set the alert before 100%?
5
15
Clean up: Delete the resource group. Why is this the most important step of any cloud demo?
5

Reflection (no points, but the best answer gets +20)

Dynatrace uses DQL, Azure uses KQL. What's one thing you could find faster with a query than by clicking through screens?

Facilitator checklist: sign in to Dynatrace and Azure before the session · pre-create the web app the day before · with 120 phones, run the app on the Basic B1 plan for the day (the Free F1 plan has a daily CPU quota and can stop mid-session) · show a QR code for the web app URL · delete rg-chaicart-demo afterwards.