Keep ChaiCart Alive
Yesterday: design your cloud startup.
Today: diagnose failures and practise recovery.
Our loop: observe → explain → restore → learn.
Azure orders and dependencies are simulated. Your Live architecture is a teaching design; it does not provision cloud resources.
Day 2 — Keep ChaiCart Alive
Morning · 09:30–12:40
- X networking icebreaker
- Reliability, monitoring & APM
- Who Killed Checkout?
- Logs, metrics & Treasure Hunt
- SRE & Error Budget Poker
Afternoon · 13:25–15:00
- Blameless postmortem
- DevOps & Digital Delivery Factory
- CI/CD demonstration
- Security, scale & Cloud Bill Shock
Final session · 15:10–17:00
- Business systems & Follow the Order
- Trends & Career Tarot
- Skills and Q&A
- Quiz, Demo Day & awards
Breaks: 11:00 & 15:00 · Lunch: 12:40–13:25 · X activity: +100 credits per team
Yesterday in one word
One team from each region shouts one word that sums up Day 1. Then, the standings.
Today's roles matter
SRE owns the error budget · CFO owns the cloud bill · CTO defends the architecture · COO keeps time and hands in forms · CEO pitches at Demo Day
Learn in public, together
- Create or update your X profile: a short bio describing what you are learning.
- Find two workshop peers and follow each other if you choose. Use the QR to find @theharithsa.
- Throughout Day 2, post workshop photos, learning moments or takeaways on X with #ChaiCartCloudWorkshop. Ask permission before sharing someone’s photo.
- Browse tagged posts and help a peer. Submit your best post link to your captain before the final awards.
Best Workshop Post or Photo: one winner at closing, judged on learning value, creativity and workshop spirit—not likes. Posting is optional.
Earn +100: every teammate completes the activity; COO reports completion to the captain. Once per team.
Already have an account? Improve it; existing connections and follows count. Signup delayed? Draft your profile now and finish during a break; ask your captain to record completion. Never share passwords or OTPs.
ChaiCart is LIVE in 5 cities! 🎉
Push notification sent to 2.1 million users…
09:04 — “Payment stuck on spinner” 😡😡😡
Performance &
Reliability
Users don't care about your CPU. They care: “Is it fast? Does it work?”
Four things that decide if they come back
Latency
How long did it take? (Look at p95/p99, not just average.)
Availability
Could they use it at all?
Errors
Did the action fail?
Throughput
How many orders per minute can we handle?
Why p95, not average? If 95 people wait 0.2s and 5 people wait 30s, the average looks “fine” — those 5 are tweeting about you. 🐦
How much downtime per year is allowed?
| Availability | Nickname | Allowed downtime / year | per month |
|---|---|---|---|
| 99% | “two nines” | 3.65 days | ~7.3 hours |
| 99.9% | “three nines” | 8.76 hours | ~43.8 minutes |
| 99.99% | “four nines” | 52.6 minutes | ~4.4 minutes |
| 99.999% | “five nines” | 5.26 minutes | ~26 seconds |
Each extra nine costs roughly 10× more effort & money. Does a chai app need five nines? Does UPI?
How reliable systems are built
🧬 Redundancy
No single point of failure. Multiple instances across availability zones.
⏱️ Timeouts
Never wait forever. Fail fast, free the resources.
🔁 Retries with backoff
Try again — but slower each time, so you don't DDoS yourself.
🔌 Circuit breaker
If payment gateway is down, stop calling it for a while.
🪂 Graceful degradation
Recommendations broken? Still sell chai — just hide the “You may also like”.
🧪 Chaos engineering
Break things on purpose (Netflix's Chaos Monkey) before they break on their own.
Monitoring, Observability
& APM
Monitoring vs Observability
| Monitoring | Observability | |
|---|---|---|
| Question | “Is it broken?” | “Why is it broken?” |
| Handles | Known problems (known unknowns) | Surprises (unknown unknowns) |
| Approach | Predefined dashboards & thresholds | Explore rich, connected telemetry |
| Analogy | Smoke alarm 🚨 | Fire investigator 🕵️ |
APM — Application Performance Management/Monitoring
Service maps · response time & failure rate per endpoint · code-level hotspots · database calls · Real User Monitoring (RUM) · synthetic checks · AI-assisted root cause. Tools: Dynatrace, Azure Application Insights, Datadog, New Relic, Grafana stack, Elastic…
The 4 Golden Signals (Google SRE)
Latency
Time to serve a request
Traffic
Demand: requests/sec, orders/min
Errors
Rate of failed requests
Saturation
How “full” is it? CPU, memory, connection pools, queues
RED method — for services
Rate · Errors · Duration
USE method — for resources
Utilization · Saturation · Errors
💡 Remember “saturation” and “connection pools”… you'll need it in 2 minutes.
Who Killed Checkout?
Open the incident evidence in ChaiCart Live. Now. 🕵️
Six suspects. One killer. Maybe an accomplice.
🗄️ Postgres Pete
The orders database. “I was just sitting there. Nobody even talked to me.”
🚀 Deploy Dave
payment-service v2.3.1, deployed 08:58. “Tiny config tweak. FinOps approved it!”
💳 PayFast Priya
External payment gateway. “My status page is green.”
💾 Disk-Full Dinesh
Log server at 97% disk. “I've been this full for two days. Nobody cares.”
⚡ Redis Rani
The cache. “Hit ratio 94%, same as always.”
📈 Traffic Tara
Launch traffic surge, 3×. “You planned for 5×. Don't blame me.”
Rules of the investigation 🕵️
- Your app evidence: Exhibit A Metrics · B Logs · C Traces · D Change log · E Witnesses
- Fill the Accusation Form: killer, weapon, accomplice, evidence, fix
- Evidence must cite at least one metric, one log line and one trace
- The COO submits the accusation in ChaiCart Live. The app records the submission time; your region captain reviews the evidence. Speed counts
Scoring 💰
- Correct killer: +200
- Correct weapon (mechanism): +100
- Correct accomplice: +50
- Good fix + prevention: +50
- First correct team in the room: +100
- Wrong killer accused: −50
Let's walk the evidence chain
- 🟢 Postgres Pete — innocent. DB CPU flat at ~28%. His active connections dropped from ~42 to 5 at 08:59. He wasn't overloaded — he was starved.
- 🟢 PayFast Priya — innocent. Latency flat ~205 ms. Failed traces never even reach her. Calls to her dropped.
- 🟢 Redis Rani & Disk-Full Dinesh — innocent. Flat for days. Classic red herrings (real issues, wrong incident).
- 🔴 Deploy Dave — THE KILLER. 08:58 deploy set
maximumPoolSize=5(was 50) — ticket FINOPS-482 “reduce DB cost”, load test skipped. - 🟠 Traffic Tara — THE ACCOMPLICE. 3× traffic at 09:00 exhausted 5 connections → requests wait 30 s →
Connection is not available→ checkout 504s.
What real SREs take away
🕐 “What changed?” is question #1
Most incidents follow a change. Put deployment events on your dashboards.
🔗 Correlate signals
Metric told us when, logs told us what, trace told us where. Alone, each was ambiguous.
🪣 Saturation hides in odd places
Not CPU, not memory — a connection pool. Watch pools, queues, threads.
💰 Cost vs reliability is a trade-off
A “FinOps win” caused the outage. Cost changes need load tests & canaries too.
Fix: roll back now. Prevent: load test in CI, canary deploys, SLO alerts on payment latency, pool-saturation alerts.
Logs, Metrics &
Distributed Tracing
What each signal looks like
📈 Metric
checkout.success_rate 08:59 99.1% 09:01 97.8% 09:03 38.2%
Cheap, aggregated, great for alerts & trends.
📓 Log
09:02:43 ERROR [payment] HikariPool-1 - Connection is not available, request timed out after 30000ms
Rich detail about one event.
🧵 Trace
checkout ███████████ 30.0s
├ cart ▏ 19ms
└ payment ██████████ 30.0s
└ getConnection ███ 30.0s
One request's journey across services.
Plus: events (deployments, config changes), profiles (code hotspots), real-user data. Standard for collecting them all: OpenTelemetry (CNCF).
Remember the kitchen order ticket? 🎫
- Trace ID = order ticket number, created at the front door
- Span = one station's work (start time, duration, status)
- Context propagation = passing the ticket number in each HTTP header (
traceparent) - The tool stitches spans into a waterfall
Hunt for clues in a real environment
Part A — Dynatrace
- Find the service with the highest failure rate
- Open a failed trace — which span is slowest?
- Open a Problem — what's the root cause?
- Query ERROR logs with DQL
Part B — Azure
- Deploy a free-tier web app
- Turn on Application Insights
- Watch Live Metrics while everyone hits the URL
- Set a cost budget alert
Answers and proof: ChaiCart Live · +10 per correct answer
Asking questions of your telemetry
// How many errors, and when did they spike?
fetch logs
| filter loglevel == "ERROR"
| makeTimeseries errors = count(), interval: 5m
// Which services are the slowest (p95)?
fetch spans
| summarize p95 = percentile(duration, 95), by: { service.name }
| sort p95 desc
| limit 10
// Is any host saturated?
timeseries cpu = avg(dt.host.cpu.usage), by: { dt.entity.host }
Observability = being able to ask new questions without shipping new code.
Site Reliability
Engineering (SRE)
“SRE is what happens when you ask a software engineer to design an operations team.”
— Ben Treynor Sloss, Google
🎯 SLOs & error budgets
Reliability as a number, not a feeling
🤖 Eliminate toil
Automate repetitive manual work
🚒 On-call & incidents
Clear roles, runbooks, calm response
📝 Blameless postmortems
Learn from failure, don't punish it
Making reliability measurable
| Meaning | ChaiCart example | |
|---|---|---|
| SLI Indicator | What you measure | % of checkout requests that succeed in < 2 s |
| SLO Objective | Your internal target | 99.9% over 30 days |
| SLA Agreement | Promise to customers, with penalties | 99.5% or we refund delivery fees |
💸 Error budget = 100% − SLO
99.9% per 30 days ⇒ ~43 minutes of “allowed badness”. Budget left? Ship features fast. Budget gone? Freeze & fix reliability.
Spend your 43 minutes wisely
- Your SRE holds the budget: 43 minutes
- I reveal an event card — each team chooses A or B in 60 seconds
- Some choices need a dice roll 🎲
- Track minutes in the app
End of the month
- Budget left: +3 credits per minute
- Budget exhausted: −100 credits & feature freeze (no more “ship” choices)
- Features shipped: credits shown on the app card
Write the postmortem for “Who Killed Checkout?”
Template
- Summary (2 lines)
- Impact — users, money, minutes
- Timeline (08:58 → recovery)
- Root cause & contributing factors
- What went well 👍
- Action items — owner + date
🚫 Blameless rules
Ask “what” and “how”, never “who”.
Blame: “Dave broke prod.”
Blameless: “Our pipeline allowed a config change to production without a load test.”
People aren't the root cause — systems that let mistakes reach production are.
DevOps, CI/CD &
Cloud-Native Development
DevOps = culture + automation + measurement
CI — Continuous Integration
Every commit is built & tested automatically.
CD — Continuous Delivery/Deployment
Every green build can go (or goes) to production.
IaC — Infrastructure as Code
Servers defined in files (Bicep, Terraform) — versioned & reviewed.
Sequential delivery vs continuous flow
Open ChaiCart Live. Six virtual changes per round: Plan → Build → Test → Deploy.
Round 1 — Serial (3 min)
The COO moves one ticket through every stage before starting the next. Observe the queue.
Round 2 — Team flow (3 min)
CEO plans, CTO builds, SRE tests, COO deploys. Different tickets move in parallel. Failed tests return to Build.
Measure in the app
- Successful deployments and frequency
- Median lead time
- Failed tests and correction time
These are teaching proxies for DORA. Choose the region winner by successful deliveries, then fewer failures, then lower lead time. No paper required.
The DORA metrics
🚀 Deployment frequency
How often you ship
⏳ Lead time for changes
Commit → production
💥 Change failure rate
% of deploys causing incidents
🩹 Time to restore
How fast you recover
Safer ways to deploy
Blue-green: two identical environments, flip traffic · Canary: 5% of users first, watch metrics, then 100% · Feature flags: ship code dark, turn on later · Automatic rollback when SLOs break
Deploy Dave would have been caught by a canary + an SLO alert. 😉
From commit to cloud in minutes
# .github/workflows/deploy.yml (GitHub Actions → Azure App Service)
on: { push: { branches: [ main ] } }
jobs:
build-and-deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 20 }
- run: npm ci && npm test # ← CI: fail fast
- uses: azure/webapps-deploy@v3 # ← CD: ship it
with:
app-name: chaicart-demo
publish-profile: ${{ secrets.AZURE_WEBAPP_PUBLISH_PROFILE }}
Show: push a change to the page title → pipeline turns green → refresh the site. Then push a failing test → pipeline blocks the deploy. 🛑
Security, Scalability
& Cost
Provider secures the cloud. You secure what's in the cloud.
| IaaS | PaaS | SaaS | |
|---|---|---|---|
| Data & access (who can see what) | You | You | You |
| Identities & accounts | You | You | You |
| Application | You | You | Provider |
| Network controls | You | Shared | Provider |
| Operating system patches | You | Provider | Provider |
| Physical hosts, network, data centre | Provider | Provider | Provider |
⚠️ Most cloud breaches are customer misconfigurations — public storage, leaked keys, over-privileged accounts, no MFA.
The cloud security starter pack 🔐
🪪 Identity is the new perimeter
MFA everywhere, least privilege, no shared admin accounts.
🔑 No secrets in code
Use a vault (Azure Key Vault) and managed identities.
🔒 Encrypt everything
In transit (TLS) and at rest.
🧱 Zero Trust
Never trust, always verify — even inside the network.
📜 Policy as code
Block public storage accounts automatically.
👀 Detect & respond
Security logs, threat detection, vulnerability scanning.
Growing without breaking
⬆️ Vertical scaling (scale up)
Bigger machine. Simple, but has a ceiling & often needs a restart.
↔️ Horizontal scaling (scale out)
More machines. Needs stateless apps. Practically limitless.
🤖 Autoscaling
Add/remove instances on CPU, queue length, schedule
⚡ Caching
Don't hit the DB for the same menu 1M times
📬 Queues
Absorb spikes, process steadily
🔥 Pre-warming
Scale before the IPL toss — autoscaling can be too slow for sudden spikes
The cloud is cheap… until the bill arrives 🧾
📏 Rightsize
That VM averaging 8% CPU? Make it smaller.
📅 Commit
Reserved instances / savings plans for steady workloads.
🎲 Spot VMs
Big discounts for interruptible batch jobs.
🌙 Schedule
Shut dev/test down at night & weekends.
🏷️ Tag & show back
Every resource tagged with team & app.
🚨 Budgets & alerts
Alert at 80% of budget, not after 300%.
FinOps = engineering + finance + business sharing responsibility for cloud spend. Also watch data egress and log/telemetry volume — silent bill killers.
The month-end bill arrives…
- Each team's CFO draws 3 Bill Shock cards
- Some cards depend on your Day 1 architecture — CTO must prove it with the frozen app design!
- Apply the credit change to the leaderboard
- For every negative card: name the control that would have prevented it → get half the loss back
Sample shocks
🎮 Intern left a GPU VM running all weekend
🪣 Storage bucket set to public — data leaked
🔑 API keys pushed to public GitHub
📜 Debug logs kept forever
Business Systems
in the Cloud
The software that actually runs a company
What runs behind every “Order placed” message
🤝 CRM
Customers, leads, loyalty, support
Salesforce, Dynamics 365, Zoho CRM
🏭 ERP
Finance, procurement, inventory, accounting
SAP S/4HANA, Oracle Fusion, Dynamics 365
🚚 SCM
Suppliers, warehouses, logistics
SAP IBP, Oracle SCM, Blue Yonder
👥 HCM
Hiring, payroll, attendance
Workday, SAP SuccessFactors, Darwinbox
🛠️ ITSM
Incidents, changes, requests
ServiceNow, Jira Service Management
📊 BI & analytics
Dashboards & decisions
Power BI, Tableau, Looker
🛒 Commerce
Storefront, catalogue, orders
Shopify, custom apps
💬 Collaboration
Email, chat, docs
Microsoft 365, Google Workspace, Slack
Almost all of these are now delivered as SaaS — that's cloud computing's biggest business impact.
One chai order, seven business systems
- 7 volunteers each open their digital business system card
- The Order Token travels from person to person
- Each system confirms the digital handoff and says what it did out loud
- Then chaos: ERP is down for maintenance — what still works? What queues up?
Integration is where the real work is
🔌 APIs
Systems talk via REST/GraphQL APIs.
📣 Events
“OrderPlaced” published once, consumed by many (Service Bus, Event Hubs, Kafka).
🧩 iPaaS
Integration platforms: Logic Apps, MuleSoft, Boomi.
✅ Why SaaS business systems win
No upgrades to run · mobile-ready · built-in AI · pay per user · fast rollout
⚠️ Watch out for
Vendor lock-in · data residency · integration sprawl · end-to-end visibility across SaaS + custom apps
“Every company is becoming a software company”
Speed to market
Launch in a new city next week, not next year.
Platform & API economy
UPI, ONDC, maps, payments — businesses built by combining APIs.
Data-driven decisions
Which city orders ginger chai at 4 PM? Cloud analytics knows.
AI everywhere
Chatbots, demand forecasting, fraud detection — rented, not built.
Scale for small players
A 5-person startup gets the same infrastructure as a giant.
Resilience = revenue
Downtime is now lost sales, not just an IT problem.
Trends, Careers
& Skills
Top trends to watch 🔮
🤖 GenAI & AI agents
GPU clouds, copilots, agentic ops
🏗️ Platform engineering
Internal developer platforms (e.g. Backstage)
💰 FinOps
Cost as an engineering metric
🏛️ Sovereign cloud
Data stays in-country
🔭 AIOps & OpenTelemetry
AI-assisted operations, open telemetry standards
⚡ Serverless
No servers to manage
📡 Edge & 5G
Compute close to devices
🛡️ Zero Trust & DevSecOps
Security built into pipelines
🌱 Green cloud
Carbon-aware computing
☁️☁️ Multi-cloud
Portability via Kubernetes
Draw your cloud destiny
🏛️ Cloud Architect
🔁 DevOps Engineer
🚒 Site Reliability Engineer
🏗️ Platform Engineer
🛡️ Cloud Security Engineer
💰 FinOps Analyst
🔭 Observability Engineer
🤖 AI / MLOps Engineer
Each member chooses a career card in My workshop. Read your “destiny”. Accept it, or swap once with a teammate. Then fill in your row of the 90-day plan in your individual app plan.
What to learn — in order
- Foundations: Linux, networking (DNS, HTTP, TCP/IP), Git
- One cloud, properly: compute, storage, networking, identity
- Scripting: Python or Bash/PowerShell
- Containers & Kubernetes
- Infrastructure as Code: Terraform / Bicep
- CI/CD: GitHub Actions / Azure DevOps
- Observability: metrics, logs, traces, SLOs
- Security basics & cost awareness
- One portfolio project on GitHub: deploy it, monitor it, write it up. It matters more than grades in interviews
- Communication, the underrated superpower
Starter certifications
Start here (2nd–4th year): Azure AZ-900 · AWS Cloud Practitioner · Google Cloud Digital Leader. Student offers such as Azure for Students give free credits to practise.
Next step: AZ-104 · AWS Solutions Architect Associate · Google Associate Cloud Engineer
Specialise: CKA/CKAD · Terraform Associate · AZ-400 · AZ-500 · FinOps Certified Practitioner · Dynatrace Associate
Free learning: Microsoft Learn · AWS Skill Builder · Google Cloud Skills Boost · Dynatrace University
Hot-seat questions 🔥
“If AI can write Terraform and read dashboards, what will a cloud engineer do in 5 years?”
“Should an Indian bank ever run core banking on public cloud?”
“Is multi-cloud a smart strategy or an expensive fashion?”
“Who should get paged at 3 AM — the developer who wrote the code or the ops team?”
“What's the most expensive cloud mistake you can imagine at ChaiCart?”
“Open floor — ask me anything about working in the industry.”
Final quiz
Same rules as yesterday: COO submits the team answer in ChaiCart Live. The facilitator locks each question and applies scores.
Demo Day
Region round: each team pitches for 90 seconds to its region. The region votes for its finalist.
Grand final: the 4 region finalists pitch for 3 minutes on stage.
- Your cloud and architecture (show the saved app design)
- What broke, and your postmortem action items
- Your SLO and remaining error budget
- Your final Cloud Credits balance
- One thing you'll do differently
Awards
- Richest Startup (overall leaderboard)
- Region Champions (top of each region)
- Sherlock Award (murder mystery)
- Best Architecture
- Best Blameless Postmortem
- Most Creative Pitch (grand final)
- Best Workshop Post or Photo (X)
You built it. You broke it.
You kept it alive. ☕☁️
Please fill in the feedback form · Save the digital cheat-sheet link · Keep learning