all projects

case study~/projects/deploysentinel-ai

DeploySentinel AI

Turns AWS deployment failures into evidence-backed root-cause analysis and human-approved automated recovery. A personal hackathon build on ECS/Fargate, Bedrock, Step Functions and Terraform in the Sydney Region.

Type
Personal project · AWS hackathon build
Region
ap-southeast-2 (Sydney)
Infrastructure
57 Terraform-managed resources
Tests
17 frontend behaviour tests (Vitest) · backend unit tests
Status
INC-001 resolved; recovery execution disabled again after the demo

~/problem

A failed deployment is loud but unexplained

ECS reports a circuit breaker, the load balancer reports unhealthy targets, and the logs hold the real reason. The on-call engineer spends the first twenty minutes stitching those together before anyone decides what to do. Rollbacks are then performed by hand, under pressure, with no record of why the decision was right.

~/solution

Evidence first, AI second, humans in charge

DeploySentinel captures the failed deployment's own evidence, asks Amazon Bedrock for a structured root-cause analysis that is validated against that evidence, shows an operator one page they can act on, and runs recovery to the known-good revision only after a human approves it.

  1. 01

    A deployment fails

    Revision 2 of the demo service shipped configuration that pointed it at a dependency it could not reach. Tasks reached RUNNING, the health endpoint returned HTTP 500 and the load balancer marked the targets unhealthy: loud, but unexplained.

  2. 02

    Evidence is captured, not guessed

    ECS deployment and task state, ALB target health, CloudWatch application logs and the task-definition diff between revisions are collected into one compact, immutable evidence package.

  3. 03

    Bedrock explains the failure

    Amazon Bedrock (Nova Pro, Converse API) returns a structured root-cause analysis. Every signal it cites must exist in the captured evidence catalogue, or the analysis is rejected; confidence is labelled an AI assessment, never certainty.

  4. 04

    An operator reads one page

    Root cause, evidence, causal chain, configuration diff, timeline and a recommended recovery to the known-good revision are stored in DynamoDB and served by API Gateway and Lambda to a static dashboard.

  5. 05

    A human approves recovery

    Review recovery, request recovery. Only then does AWS Step Functions run a least-privilege recovery Lambda that rolls the service back to the known-good revision.

  6. 06

    Resolved only after verification

    The incident is marked resolved only after independent checks of ECS, the load balancer and the application itself pass; the timeline records each step.

~/architecture

How the pieces fit

A demo service runs on ECS/Fargate behind an Application Load Balancer. Its deployment events, target health and CloudWatch logs feed the evidence collection, which also diffs the task definitions. Amazon Bedrock produces the grounded analysis; API Gateway and Lambda serve the incident, analysis and timeline from DynamoDB to a static dashboard on Amplify Hosting. A human approval from that dashboard starts an AWS Step Functions execution that drives a least-privilege recovery Lambda, then verifies ECS, the load balancer and the application before the incident is resolved.

The project's own architecture, as deployed and documented in its repository. It is not an employer's system.

  • Amazon ECS on AWS Fargate
  • Application Load Balancer
  • Amazon ECR
  • AWS Lambda
  • Amazon API Gateway (HTTP API)
  • Amazon DynamoDB
  • Amazon Bedrock (Nova Pro)
  • AWS Step Functions
  • Amazon CloudWatch
  • AWS Amplify Hosting
  • AWS IAM
  • Terraform
  • React + Vite dashboard
  • Python (Lambdas, tooling)

~/ai-analysis

What the AI is allowed to say

The model never sees raw dumps. The failed deployment's evidence is packaged into a compact, immutable catalogue of observed signals, and Amazon Bedrock (Nova Pro, through the Converse API) is asked for a structured root-cause analysis against it. Every signal the analysis cites must exist in that catalogue or the result is rejected, and the confidence it reports is presented as an AI assessment, not a measurement. The recommendation is always a recovery to the known-good revision, never an edit the model invents.

~/reliability-and-security

Guardrails

  • Recovery never runs on its own: a human approves it from the dashboard, and execution is disabled again outside the demo window.
  • The approval request carries an empty body: the browser never names a cluster, service or revision. The recovery Lambda is least-privilege.
  • No IAM users or long-lived access keys: temporary credentials only.
  • Infrastructure changes are phase-gated: a saved Terraform plan, a hash-pinned machine review that asserts the exact change set, then apply. Evidence for every phase is committed under docs/.
  • The dashboard is static hosting with a SPA rewrite and strict security headers; the backend is API Gateway, Lambda, DynamoDB and Step Functions.
  • Ship-gate scripts check the public dashboard, API reads with CORS, denied origins and the recovery state machine before anything is called done.

~/results

What was verified

  • A real ECS deployment failure (revision 2) was investigated end to end with Amazon Bedrock.
  • A human-approved recovery rolled the service back to the known-good revision 1 through Step Functions.
  • ECS, the load balancer and the application were verified before the incident was resolved.
  • Terraform matches the recovered infrastructure: plan shows no changes.

Scope: a controlled, reversible failure in the project's own demo service. This is a hackathon build, not a production product, and no adoption or benchmark figures are claimed.