Cloud Computing & Business Systems Workshop · Day 2

Keep ChaiCart Alive

Build the system on Day 1, test a blocked flow on Day 2, then recover.

Yesterday: design your cloud startup.

Today: diagnose failures and practise recovery.

Our loop: observe → explain → restore → learn.

Azure orders and dependencies are simulated. Your Live architecture is a teaching design; it does not provision cloud resources.

Today’s agenda · 09:30–17:00

Day 2 — Keep ChaiCart Alive

Morning · 09:30–12:40

  • X networking icebreaker
  • Reliability, monitoring & APM
  • Who Killed Checkout?
  • Logs, metrics & Treasure Hunt
  • SRE & Error Budget Poker

Afternoon · 13:25–15:00

  • Blameless postmortem
  • DevOps & Digital Delivery Factory
  • CI/CD demonstration
  • Security, scale & Cloud Bill Shock

Final session · 15:10–17:00

  • Business systems & Follow the Order
  • Trends & Career Tarot
  • Skills and Q&A
  • Quiz, Demo Day & awards

Breaks: 11:00 & 15:00 · Lunch: 12:40–13:25 · X activity: +100 credits per team

Warm-up · 3 min

Yesterday in one word

One team from each region shouts one word that sums up Day 1. Then, the standings.

Open leaderboard

Today's roles matter

SRE owns the error budget · CFO owns the cloud bill · CTO defends the architecture · COO keeps time and hands in forms · CEO pitches at Demo Day

NETWORKING 3/3 · 5 min kickoff + all day · +100 / team

Learn in public, together

  1. Create or update your X profile: a short bio describing what you are learning.
  2. Find two workshop peers and follow each other if you choose. Use the QR to find @theharithsa.
  3. Throughout Day 2, post workshop photos, learning moments or takeaways on X with #ChaiCartCloudWorkshop. Ask permission before sharing someone’s photo.
  4. Browse tagged posts and help a peer. Submit your best post link to your captain before the final awards.

Best Workshop Post or Photo: one winner at closing, judged on learning value, creativity and workshop spirit—not likes. Posting is optional.

QR code for Vishruth’s X profile

X · @theharithsa

x.com/theharithsa

Scan to open the presenter’s profile.

Earn +100: every teammate completes the activity; COO reports completion to the captain. Once per team.

Already have an account? Improve it; existing connections and follows count. Signup delayed? Draft your profile now and finish during a break; ask your captain to record completion. Never share passwords or OTPs.

● BREAKING NEWS · 09:00 AM

ChaiCart is LIVE in 5 cities! 🎉

Push notification sent to 2.1 million users…

09:04 — “Payment stuck on spinner” 😡😡😡

CHAPTER 1

Performance &
Reliability

Users don't care about your CPU. They care: “Is it fast? Does it work?”

What users actually feel

Four things that decide if they come back

🐢

Latency

How long did it take? (Look at p95/p99, not just average.)

✅

Availability

Could they use it at all?

💥

Errors

Did the action fail?

🚦

Throughput

How many orders per minute can we handle?

Why p95, not average? If 95 people wait 0.2s and 5 people wait 30s, the average looks “fine” — those 5 are tweeting about you. 🐦

GAME · Guess the Downtime ⏳ · closest team +20 each round

How much downtime per year is allowed?

AvailabilityNicknameAllowed downtime / yearper month
99%“two nines”3.65 days~7.3 hours
99.9%“three nines”8.76 hours~43.8 minutes
99.99%“four nines”52.6 minutes~4.4 minutes
99.999%“five nines”5.26 minutes~26 seconds

Each extra nine costs roughly 10× more effort & money. Does a chai app need five nines? Does UPI?

Reliability patterns

How reliable systems are built

🧬 Redundancy

No single point of failure. Multiple instances across availability zones.

⏱️ Timeouts

Never wait forever. Fail fast, free the resources.

🔁 Retries with backoff

Try again — but slower each time, so you don't DDoS yourself.

🔌 Circuit breaker

If payment gateway is down, stop calling it for a while.

🪂 Graceful degradation

Recommendations broken? Still sell chai — just hide the “You may also like”.

🧪 Chaos engineering

Break things on purpose (Netflix's Chaos Monkey) before they break on their own.

CHAPTER 2

Monitoring, Observability
& APM

Not the same thing

Monitoring vs Observability

MonitoringObservability
Question“Is it broken?”“Why is it broken?”
HandlesKnown problems (known unknowns)Surprises (unknown unknowns)
ApproachPredefined dashboards & thresholdsExplore rich, connected telemetry
AnalogySmoke alarm 🚨Fire investigator 🕵️

APM — Application Performance Management/Monitoring

Service maps · response time & failure rate per endpoint · code-level hotspots · database calls · Real User Monitoring (RUM) · synthetic checks · AI-assisted root cause. Tools: Dynatrace, Azure Application Insights, Datadog, New Relic, Grafana stack, Elastic…

What to watch

The 4 Golden Signals (Google SRE)

⏱️

Latency

Time to serve a request

🚗

Traffic

Demand: requests/sec, orders/min

❌

Errors

Rate of failed requests

🪣

Saturation

How “full” is it? CPU, memory, connection pools, queues

RED method — for services

Rate · Errors · Duration

USE method — for resources

Utilization · Saturation · Errors

💡 Remember “saturation” and “connection pools”… you'll need it in 2 minutes.

🔪 Murder Mystery

Who Killed Checkout?

Open the incident evidence in ChaiCart Live. Now. 🕵️

The suspects

Six suspects. One killer. Maybe an accomplice.

🗄️ Postgres Pete

The orders database. “I was just sitting there. Nobody even talked to me.”

🚀 Deploy Dave

payment-service v2.3.1, deployed 08:58. “Tiny config tweak. FinOps approved it!”

💳 PayFast Priya

External payment gateway. “My status page is green.”

💾 Disk-Full Dinesh

Log server at 97% disk. “I've been this full for two days. Nobody cares.”

⚡ Redis Rani

The cache. “Hit ratio 94%, same as always.”

📈 Traffic Tara

Launch traffic surge, 3×. “You planned for 5×. Don't blame me.”

INVESTIGATION · 25 min

Rules of the investigation 🕵️

  • Your app evidence: Exhibit A Metrics · B Logs · C Traces · D Change log · E Witnesses
  • Fill the Accusation Form: killer, weapon, accomplice, evidence, fix
  • Evidence must cite at least one metric, one log line and one trace
  • The COO submits the accusation in ChaiCart Live. The app records the submission time; your region captain reviews the evidence. Speed counts

Scoring 💰

  • Correct killer: +200
  • Correct weapon (mechanism): +100
  • Correct accomplice: +50
  • Good fix + prevention: +50
  • First correct team in the room: +100
  • Wrong killer accused: −50
The reveal 🎭

Let's walk the evidence chain

  1. 🟢 Postgres Pete — innocent. DB CPU flat at ~28%. His active connections dropped from ~42 to 5 at 08:59. He wasn't overloaded — he was starved.
  2. 🟢 PayFast Priya — innocent. Latency flat ~205 ms. Failed traces never even reach her. Calls to her dropped.
  3. 🟢 Redis Rani & Disk-Full Dinesh — innocent. Flat for days. Classic red herrings (real issues, wrong incident).
  4. 🔴 Deploy Dave — THE KILLER. 08:58 deploy set maximumPoolSize=5 (was 50) — ticket FINOPS-482 “reduce DB cost”, load test skipped.
  5. 🟠 Traffic Tara — THE ACCOMPLICE. 3× traffic at 09:00 exhausted 5 connections → requests wait 30 s → Connection is not available → checkout 504s.
Lessons from the crime scene

What real SREs take away

🕐 “What changed?” is question #1

Most incidents follow a change. Put deployment events on your dashboards.

🔗 Correlate signals

Metric told us when, logs told us what, trace told us where. Alone, each was ambiguous.

🪣 Saturation hides in odd places

Not CPU, not memory — a connection pool. Watch pools, queues, threads.

💰 Cost vs reliability is a trade-off

A “FinOps win” caused the outage. Cost changes need load tests & canaries too.

Fix: roll back now. Prevent: load test in CI, canary deploys, SLO alerts on payment latency, pool-saturation alerts.

CHAPTER 3

Logs, Metrics &
Distributed Tracing

The three pillars (and friends)

What each signal looks like

📈 Metric

checkout.success_rate
08:59  99.1%
09:01  97.8%
09:03  38.2%

Cheap, aggregated, great for alerts & trends.

📓 Log

09:02:43 ERROR [payment]
HikariPool-1 - Connection
is not available, request
timed out after 30000ms

Rich detail about one event.

🧵 Trace

checkout   ███████████ 30.0s
 ├ cart    ▏ 19ms
 └ payment ██████████ 30.0s
    └ getConnection ███ 30.0s

One request's journey across services.

Plus: events (deployments, config changes), profiles (code hotspots), real-user data. Standard for collecting them all: OpenTelemetry (CNCF).

How tracing works

Remember the kitchen order ticket? 🎫

  • Trace ID = order ticket number, created at the front door
  • Span = one station's work (start time, duration, status)
  • Context propagation = passing the ticket number in each HTTP header (traceparent)
  • The tool stitches spans into a waterfall
📱 App — trace 4bf92f…
checkout-service · span 1
cart-service · span 2
payment-service · span 3
orders-db · span 4
PayFast API · span 5
HANDS-ON · Observability Treasure Hunt 🗺️ · 35 min

Hunt for clues in a real environment

Part A — Dynatrace

  • Find the service with the highest failure rate
  • Open a failed trace — which span is slowest?
  • Open a Problem — what's the root cause?
  • Query ERROR logs with DQL

Part B — Azure

  • Deploy a free-tier web app
  • Turn on Application Insights
  • Watch Live Metrics while everyone hits the URL
  • Set a cost budget alert

Answers and proof: ChaiCart Live · +10 per correct answer

LIVE DEMO · DQL

Asking questions of your telemetry

// How many errors, and when did they spike?
fetch logs
| filter loglevel == "ERROR"
| makeTimeseries errors = count(), interval: 5m

// Which services are the slowest (p95)?
fetch spans
| summarize p95 = percentile(duration, 95), by: { service.name }
| sort p95 desc
| limit 10

// Is any host saturated?
timeseries cpu = avg(dt.host.cpu.usage), by: { dt.entity.host }

Observability = being able to ask new questions without shipping new code.

CHAPTER 4

Site Reliability
Engineering (SRE)

What is SRE?
“SRE is what happens when you ask a software engineer to design an operations team.”
— Ben Treynor Sloss, Google

🎯 SLOs & error budgets

Reliability as a number, not a feeling

🤖 Eliminate toil

Automate repetitive manual work

🚒 On-call & incidents

Clear roles, runbooks, calm response

📝 Blameless postmortems

Learn from failure, don't punish it

The SLI → SLO → SLA ladder

Making reliability measurable

MeaningChaiCart example
SLI IndicatorWhat you measure% of checkout requests that succeed in < 2 s
SLO ObjectiveYour internal target99.9% over 30 days
SLA AgreementPromise to customers, with penalties99.5% or we refund delivery fees

💸 Error budget = 100% − SLO

99.9% per 30 days ⇒ ~43 minutes of “allowed badness”. Budget left? Ship features fast. Budget gone? Freeze & fix reliability.

GAME · Error Budget Poker 🃏 · 20 min

Spend your 43 minutes wisely

  • Your SRE holds the budget: 43 minutes
  • I reveal an event card — each team chooses A or B in 60 seconds
  • Some choices need a dice roll 🎲
  • Track minutes in the app

End of the month

  • Budget left: +3 credits per minute
  • Budget exhausted: −100 credits & feature freeze (no more “ship” choices)
  • Features shipped: credits shown on the app card
ROLE-PLAY · Blameless Postmortem · 12 min · Region best +40 · Overall best +75

Write the postmortem for “Who Killed Checkout?”

Template

  1. Summary (2 lines)
  2. Impact — users, money, minutes
  3. Timeline (08:58 → recovery)
  4. Root cause & contributing factors
  5. What went well 👍
  6. Action items — owner + date

🚫 Blameless rules

Ask “what” and “how”, never “who”.

Blame: “Dave broke prod.”

Blameless: “Our pipeline allowed a config change to production without a load test.”

People aren't the root cause — systems that let mistakes reach production are.

CHAPTER 5

DevOps, CI/CD &
Cloud-Native Development

The infinite loop ♾️

DevOps = culture + automation + measurement

📝 Plan
➜
💻 Code
➜
🔨 Build
➜
🧪 Test
➜
📦 Release
➜
🚀 Deploy
➜
⚙️ Operate
➜
🔭 Monitor
↺

CI — Continuous Integration

Every commit is built & tested automatically.

CD — Continuous Delivery/Deployment

Every green build can go (or goes) to production.

IaC — Infrastructure as Code

Servers defined in files (Bicep, Terraform) — versioned & reviewed.

GAME · Digital Delivery Factory · 20 min · Region best +50

Sequential delivery vs continuous flow

Open ChaiCart Live. Six virtual changes per round: Plan → Build → Test → Deploy.

Round 1 — Serial (3 min)

The COO moves one ticket through every stage before starting the next. Observe the queue.

Round 2 — Team flow (3 min)

CEO plans, CTO builds, SRE tests, COO deploys. Different tickets move in parallel. Failed tests return to Build.

Measure in the app

  • Successful deployments and frequency
  • Median lead time
  • Failed tests and correction time

These are teaching proxies for DORA. Choose the region winner by successful deliveries, then fewer failures, then lower lead time. No paper required.

Measuring DevOps

The DORA metrics

🚀 Deployment frequency

How often you ship

⏳ Lead time for changes

Commit → production

💥 Change failure rate

% of deploys causing incidents

🩹 Time to restore

How fast you recover

Safer ways to deploy

Blue-green: two identical environments, flip traffic · Canary: 5% of users first, watch metrics, then 100% · Feature flags: ship code dark, turn on later · Automatic rollback when SLOs break

Deploy Dave would have been caught by a canary + an SLO alert. 😉

LIVE DEMO · CI/CD to Azure

From commit to cloud in minutes

# .github/workflows/deploy.yml  (GitHub Actions → Azure App Service)
on: { push: { branches: [ main ] } }
jobs:
  build-and-deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 20 }
      - run: npm ci && npm test          # ← CI: fail fast
      - uses: azure/webapps-deploy@v3     # ← CD: ship it
        with:
          app-name: chaicart-demo
          publish-profile: ${{ secrets.AZURE_WEBAPP_PUBLISH_PROFILE }}

Show: push a change to the page title → pipeline turns green → refresh the site. Then push a failing test → pipeline blocks the deploy. 🛑

CHAPTER 6

Security, Scalability
& Cost

Shared responsibility model

Provider secures the cloud. You secure what's in the cloud.

IaaSPaaSSaaS
Data & access (who can see what)YouYouYou
Identities & accountsYouYouYou
ApplicationYouYouProvider
Network controlsYouSharedProvider
Operating system patchesYouProviderProvider
Physical hosts, network, data centreProviderProviderProvider

⚠️ Most cloud breaches are customer misconfigurations — public storage, leaked keys, over-privileged accounts, no MFA.

Security essentials

The cloud security starter pack 🔐

🪪 Identity is the new perimeter

MFA everywhere, least privilege, no shared admin accounts.

🔑 No secrets in code

Use a vault (Azure Key Vault) and managed identities.

🔒 Encrypt everything

In transit (TLS) and at rest.

🧱 Zero Trust

Never trust, always verify — even inside the network.

📜 Policy as code

Block public storage accounts automatically.

👀 Detect & respond

Security logs, threat detection, vulnerability scanning.

Scalability

Growing without breaking

⬆️ Vertical scaling (scale up)

Bigger machine. Simple, but has a ceiling & often needs a restart.

↔️ Horizontal scaling (scale out)

More machines. Needs stateless apps. Practically limitless.

🤖 Autoscaling

Add/remove instances on CPU, queue length, schedule

⚡ Caching

Don't hit the DB for the same menu 1M times

📬 Queues

Absorb spikes, process steadily

🔥 Pre-warming

Scale before the IPL toss — autoscaling can be too slow for sudden spikes

Cost & FinOps

The cloud is cheap… until the bill arrives 🧾

📏 Rightsize

That VM averaging 8% CPU? Make it smaller.

📅 Commit

Reserved instances / savings plans for steady workloads.

🎲 Spot VMs

Big discounts for interruptible batch jobs.

🌙 Schedule

Shut dev/test down at night & weekends.

🏷️ Tag & show back

Every resource tagged with team & app.

🚨 Budgets & alerts

Alert at 80% of budget, not after 300%.

FinOps = engineering + finance + business sharing responsibility for cloud spend. Also watch data egress and log/telemetry volume — silent bill killers.

GAME · Cloud Bill Shock 💸 · 15 min

The month-end bill arrives…

  • Each team's CFO draws 3 Bill Shock cards
  • Some cards depend on your Day 1 architecture — CTO must prove it with the frozen app design!
  • Apply the credit change to the leaderboard
  • For every negative card: name the control that would have prevented it → get half the loss back

Sample shocks

🎮 Intern left a GPU VM running all weekend
🪣 Storage bucket set to public — data leaked
🔑 API keys pushed to public GitHub
📜 Debug logs kept forever

CHAPTER 7

Business Systems
in the Cloud

The software that actually runs a company

The business systems map

What runs behind every “Order placed” message

🤝 CRM

Customers, leads, loyalty, support
Salesforce, Dynamics 365, Zoho CRM

🏭 ERP

Finance, procurement, inventory, accounting
SAP S/4HANA, Oracle Fusion, Dynamics 365

🚚 SCM

Suppliers, warehouses, logistics
SAP IBP, Oracle SCM, Blue Yonder

👥 HCM

Hiring, payroll, attendance
Workday, SAP SuccessFactors, Darwinbox

🛠️ ITSM

Incidents, changes, requests
ServiceNow, Jira Service Management

📊 BI & analytics

Dashboards & decisions
Power BI, Tableau, Looker

🛒 Commerce

Storefront, catalogue, orders
Shopify, custom apps

💬 Collaboration

Email, chat, docs
Microsoft 365, Google Workspace, Slack

Almost all of these are now delivered as SaaS — that's cloud computing's biggest business impact.

ACTIVITY · Follow the Order 📦 · 15 min

One chai order, seven business systems

  • 7 volunteers each open their digital business system card
  • The Order Token travels from person to person
  • Each system confirms the digital handoff and says what it did out loud
  • Then chaos: ERP is down for maintenance — what still works? What queues up?
🛒 Commerce — order placed
🤝 CRM — loyalty points
🏭 ERP — invoice & GST
🚚 SCM — reduce stock, reorder milk
👥 HCM — rider incentive
📊 BI — dashboard updated
🛠️ ITSM — (only if something breaks!)
Connecting the systems

Integration is where the real work is

🔌 APIs

Systems talk via REST/GraphQL APIs.

📣 Events

“OrderPlaced” published once, consumed by many (Service Bus, Event Hubs, Kafka).

🧩 iPaaS

Integration platforms: Logic Apps, MuleSoft, Boomi.

✅ Why SaaS business systems win

No upgrades to run · mobile-ready · built-in AI · pay per user · fast rollout

⚠️ Watch out for

Vendor lock-in · data residency · integration sprawl · end-to-end visibility across SaaS + custom apps

The role of cloud in digital business

“Every company is becoming a software company”

⚡

Speed to market

Launch in a new city next week, not next year.

🔗

Platform & API economy

UPI, ONDC, maps, payments — businesses built by combining APIs.

📊

Data-driven decisions

Which city orders ginger chai at 4 PM? Cloud analytics knows.

🤖

AI everywhere

Chatbots, demand forecasting, fraud detection — rented, not built.

🌍

Scale for small players

A 5-person startup gets the same infrastructure as a giant.

🔁

Resilience = revenue

Downtime is now lost sales, not just an IT problem.

CHAPTER 8

Trends, Careers
& Skills

Where the industry is heading

Top trends to watch 🔮

🤖 GenAI & AI agents

GPU clouds, copilots, agentic ops

🏗️ Platform engineering

Internal developer platforms (e.g. Backstage)

💰 FinOps

Cost as an engineering metric

🏛️ Sovereign cloud

Data stays in-country

🔭 AIOps & OpenTelemetry

AI-assisted operations, open telemetry standards

⚡ Serverless

No servers to manage

📡 Edge & 5G

Compute close to devices

🛡️ Zero Trust & DevSecOps

Security built into pipelines

🌱 Green cloud

Carbon-aware computing

☁️☁️ Multi-cloud

Portability via Kubernetes

ACTIVITY · Career Tarot 🔮 · 15 min

Draw your cloud destiny

🏛️ Cloud Architect

🔁 DevOps Engineer

🚒 Site Reliability Engineer

🏗️ Platform Engineer

🛡️ Cloud Security Engineer

💰 FinOps Analyst

🔭 Observability Engineer

🤖 AI / MLOps Engineer

Each member chooses a career card in My workshop. Read your “destiny”. Accept it, or swap once with a teammate. Then fill in your row of the 90-day plan in your individual app plan.

The skills stack

What to learn — in order

  1. Foundations: Linux, networking (DNS, HTTP, TCP/IP), Git
  2. One cloud, properly: compute, storage, networking, identity
  3. Scripting: Python or Bash/PowerShell
  4. Containers & Kubernetes
  5. Infrastructure as Code: Terraform / Bicep
  6. CI/CD: GitHub Actions / Azure DevOps
  7. Observability: metrics, logs, traces, SLOs
  8. Security basics & cost awareness
  9. One portfolio project on GitHub: deploy it, monitor it, write it up. It matters more than grades in interviews
  10. Communication, the underrated superpower

Starter certifications

Start here (2nd–4th year): Azure AZ-900 · AWS Cloud Practitioner · Google Cloud Digital Leader. Student offers such as Azure for Students give free credits to practise.

Next step: AZ-104 · AWS Solutions Architect Associate · Google Associate Cloud Engineer

Specialise: CKA/CKAD · Terraform Associate · AZ-400 · AZ-500 · FinOps Certified Practitioner · Dynatrace Associate

Free learning: Microsoft Learn · AWS Skill Builder · Google Cloud Skills Boost · Dynatrace University

Interactive discussion

Hot-seat questions 🔥

“If AI can write Terraform and read dashboards, what will a cloud engineer do in 5 years?”

“Should an Indian bank ever run core banking on public cloud?”

“Is multi-cloud a smart strategy or an expensive fashion?”

“Who should get paged at 3 AM — the developer who wrote the code or the ops team?”

“What's the most expensive cloud mistake you can imagine at ChaiCart?”

“Open floor — ask me anything about working in the industry.”

QUIZ · Day 2 recap · +10 per correct answer

Final quiz

Same rules as yesterday: COO submits the team answer in ChaiCart Live. The facilitator locks each question and applies scores.

Open live quiz controls

FINALE · ChaiCart Demo Day · Region round 15 min · Grand final 15 min

Demo Day

Region round: each team pitches for 90 seconds to its region. The region votes for its finalist.
Grand final: the 4 region finalists pitch for 3 minutes on stage.

  1. Your cloud and architecture (show the saved app design)
  2. What broke, and your postmortem action items
  3. Your SLO and remaining error budget
  4. Your final Cloud Credits balance
  5. One thing you'll do differently

Awards

  • Richest Startup (overall leaderboard)
  • Region Champions (top of each region)
  • Sherlock Award (murder mystery)
  • Best Architecture
  • Best Blameless Postmortem
  • Most Creative Pitch (grand final)
  • Best Workshop Post or Photo (X)
Thank you

You built it. You broke it.
You kept it alive. ☕☁️

Please fill in the feedback form · Save the digital cheat-sheet link · Keep learning