Checking live state…

EG334S · HyEnt Multi-Cloud Project

Three cloud cells joined by three Site-to-Site VPNs, with a two-AZ Ireland application entry and controlled recovery

Deployment environments · current and planned

Personal AWS · integrated live hub LIVE

This is the existing Ireland hub infrastructure, not a separate website or VPC. Its internet-facing Application Load Balancer spans eu-west-1a and eu-west-1b and sends normal traffic to the existing hub EC2 instance at 10.1.1.61. The cost-controlled cold-recovery Auto Scaling group is attached to the same target group and remains at desired capacity zero until recovery is triggered.

Open live AWS hub

Personal AWSTEST BED · LIVE · existing Ireland hub · external HTTP 200
School AWSDEMO TARGET · PENDING · access and deployment not started
Load balancingPASS · existing primary hub target healthy · ALB spans two Availability Zones
Hub recoveryREADY · cold-recovery ASG attached · desired capacity 0
VPN #1available · one AWS tunnel endpoint UP
Last verified30 Jul 2026, 01:59 SGT

Claim boundary. This proves the existing hub is healthy through the public ALB, the ALB spans two Availability Zones, the cold-recovery ASG is attached, and all three private paths have been exercised. It does not claim two normally running hub servers: recovery capacity intentionally remains at zero for cost control. On 30 Jul, VPN #2 and direct VPN #3 were live, and a fresh recorded integrated-ALB recovery drill restored service in 218 seconds. This is application recoverability, not full regional HA.

Decommission deadline
--
Sun 30 Aug 2026, 23:59 SGT · official penalty applies to AWS resources left running
Project week
1 / 5
Plan & set up (ahead: AWS build already live)
Milestones done
4 / 6
M1 + M2 + M3 + M4 · the system works
Programme readiness
AMBER
Engineering proven · people, approval, demo and closure gates remain

Class calendar · official 15-session timetable

All dates are SGT. Sessions 1 and 2 are Sync-Physical; sessions 3 to 15 are Sync-Online. Session 8 on 11 Aug is the National Day public-holiday make-up lesson. The VPN #2 rebuild on 25 Aug is a separate internal action so the 26 Aug assessment can run on live infrastructure. Teardown must be verified in us-east-1, eu-west-1, Azure and Cloudflare.

1What we are judged onThe standard, first. Nothing further down this page can be judged without it.

Assessment · official weighting (Project Information §4)

Checkpoint
30%
Individual · problem-solving, ownership, attendance
Presentation
30%
Individual · delivery, slides, Q&A
System demo
20%
Technical / team · completeness & troubleshooting
Report
20%
Group · neatness, clarity, logical flow

60% is individual (Checkpoint + Presentation). Build the system as a team; own and defend your own section alone.

Presentation rubric · what assessors score (individual, 30%)

Performance indicatorExcellent (top band)What it takes
Organisation & Content (60)
Organisation & supporting materials; Content
Agenda exists, coherent & interesting sequence; supporting materials used innovatively & explained in context; can explain all details, constraints and work-arounds. Have a clear agenda; every diagram/screenshot explained in context; be able to explain the whole project's details, limits and how you worked around them.
Presentation Skills (40)
Delivery; Q&A
Interesting, eloquent, enthusiastic delivery (not heavily scripted); handles all Q&A well and anticipates questions. Rehearse so you're not reading slides; pre-empt likely assessor questions and prepare answers.

Report rubric · what assessors score (group, 20%)

Performance indicatorExcellent (top band)What it takes
Report Presentation (20)
Writing; Presentation & supporting materials
Exceptionally clear, precise, concise English; few typos; professional layout; all illustrations well formatted. Proofread hard; consistent styles/margins; clean, labelled figures and screenshots.
Technical Content (30)
Organisation & structure; Literature survey; Quality of analysis
Structure entirely correct, all sections placed; exemplary range of references; well-informed, authoritative discussion of a complex problem with depth. Use a logical, complete structure; cite an exemplary range of authoritative references; show reasoned technical analysis, not just description.

Verified on 27 Jul 2026 against the supplied official Presentation Assessment Rubrics and Report Assessment Rubrics PDFs. Presentation is marked out of 100 (60 + 40); the report rubric has 50 rubric marks (20 + 30) and the report contributes 20% of the module assessment.

Report requirements & structure (Project Information §5)

Working section-order proposal PROVISIONAL

    Not yet confirmed by the team. The brief labels its contents page as a sample for reference only, so the order and owner chips shown here are planning placeholders. Use the matrix to develop the proposal; do not treat it as the approved Responsibility Assignment Table. Changes save to the shared planning matrix that every visitor sees.

    Must include

    • start Responsibility Assignment Table (name, admin no., role, topic, section, remarks)
    • evidence Diagrams, screenshots, YAML files & configuration
    • refs Bibliography of all references cited (incl. URLs)
    • format Cover page + content page
    • submit Softcopy uploaded to PoliteMall by due date

    Scope & tasks (Project Information §3) · built with Amazon Q → CloudFormation YAML

    Networks & compute

      IPSec site-to-site VPN

        Bonus marks (only after M4 · core working)

        High availability & load balancing of web servers Auto-scaling + simulated load test CloudWatch + CloudTrail + VPC Flow Logs (Bernard's slice)
        2Where we standWhere we are against that standard, and the system it refers to.

        Live multi-cloud overview · personal test bed

        reading state…
        VPN #1 · ~5,500 km · ~69 ms RTT VPN #3 DIRECT · ~15,500 km · ~225 ms RTT VPN #2 · ~11,200 km · ~189 ms RTT On-prem DC us-east-1 · N. Virginia 10.0.0.0/16 AWS public · HUB eu-west-1 · Ireland 10.1.0.0/16 ALB → existing hub EC2 10.1.1.61 cold recovery ASG · desired 0 Azure VNet southeastasia · Singapore · local Azure region 10.2.0.0/16 AWS EU-WEST-1 · EUROPE - IRELAND · DEPLOYED 2-AZ APPLICATION ENTRY APPLICATION LOAD BALANCER · ACTIVE · HTTP/80 eg334s-personal-targets · primary target healthy eu-west-1a · ACTIVE Public subnet 10.1.1.0/24 · ALB node 108.131.189.140 Hub EC2 10.1.1.61 · target HEALTHY eu-west-1b · RECOVERY READY Recovery subnet 10.1.2.0/24 · ALB node 54.216.79.2 Cold recovery ASG · desired 0 · same target group Three site-to-site VPNs: VPN1 Virginia–Ireland, VPN2 Ireland–Singapore, and direct VPN3 Virginia–Singapore. Distances are approximate; RTT is measured.

        Live status · all three private paths exercised on 30 Jul

        VPN #1 + #2 + #3 ✓ PROVEN

        VPN #1 · on-prem ↔ AWS public

        TunnelUP · 54.77.211.189
        Ping 10.0.x → 10.1.1.610% loss, ttl 254, ~69 ms
        DesignstrongSwan CGW ↔ AWS VGW, IKEv1
        PathN. Virginia ↔ Ireland (trans-Atlantic)
        Alarm
        Hub route10.0.0.0/16 present in the eu-west-1 route table, learned by propagation, not typed

        VPN #2 · AWS public ↔ Azure

        ConnectionConnected on 30 Jul · Azure gateway and connection live during the recording window
        Data planePASS · 0% loss, ~189 ms, and HTTP 200 from Azure to the Ireland hub
        DesignAzure VpnGw1AZ ↔ AWS VGW, IKEv2
        PathIreland ↔ Singapore (half the world)
        Alarm
        Direct pathVPN #3 PASS · Virginia ↔ Azure, 0% loss in both directions, ~229–230 ms
        Application-layer proof
        VPN #2 only
        curl http://10.1.1.61 from the Azure VM returned the AWS web server page, now served by the template itself · a real workload across clouds. VPN #1 was proven with ICMP, not HTTP, so this row does not apply to it.
        Hub routingeu-west-1 carries 10.0.0.0/16 from VPN #1 and 10.2.0.0/16 from VPN #2. Direct VPN #3 provides the explicit Virginia↔Azure path without relying on VGW transitivity.

        Known by design: on-prem cannot reach Azure directly (tested: 100% loss). An AWS VGW does not do transitive routing between two VPN connections. Topology is hub-and-spoke with AWS public as the hub, matching the brief. Full mesh would need a Transit Gateway.

        On-prem DC
        AWS VPC 1 · us-east-1
        N. Virginia · 10.0.0.0/16
        HUB
        AWS public
        AWS VPC 2 · eu-west-1
        Ireland · 10.1.0.0/16
        Azure VNet
        Southeast Asia (Singapore)
        10.2.0.0/16
        Reading live state…

        AWS public cloud is the hub: its single VGW terminates VPN #1 and VPN #2, each with two AWS-provided tunnel endpoints. Latencies are measured round-trip times from the genuine ICMP tests. The VGW does not perform transitive routing, so direct VPN #3 supplies the explicit Virginia↔Azure path without pretending the hub is transitive. Teardown must clear us-east-1, eu-west-1 and Azure separately. Dated state: on 30 Jul, VPN #1 and VPN #2 were active and both Azure VPN #2 and direct VPN #3 reported Connected. All three paths passed private traffic tests; the live banner and map remain authoritative if the environment changes later.

        Deployed inventory · query-level completeness

        Loading inventory…

        Nothing on this list is typed by hand. The live-data workflow runs scripts/inventory.sh, which only ever calls describe and list, and refreshes this feed automatically. Every required query records success or failure; a failed table is shown as unknown, never none. Anything here that is not inside a CloudFormation stack or the Azure resource group has to be deleted by hand at teardown, because delete-stack and az group delete will not touch it.

        3The plan to get thereCritical path, milestones, schedule and who owns what.

        Critical path · protect this chain, cut everything else first

        Green = done. Amber = AWS side done, Azure pending. Grey = not started. Anything off this chain (bonus, report polish, slides) is parallel work you sacrifice first if time runs short.

        Technical milestones · evidence-backed status

          This is a technical milestone count, not an overall completion percentage. Status is defined in the shared page source and changes only with reviewed evidence.

          5-week plan · back-planned from teardown

          Week 1
          24 Jul – 2 Aug
          Plan & set up · roles + CIDR + AWS/Azure/Q access. CIDR done AWS build + VPN #1 done (ahead)
          Week 2
          3 – 9 Aug
          Network + compute · CloudFormation for 3 networks. Deploy VPCs + VNet. Web + private servers.
          Week 3
          10 – 16 Aug
          VPN #1: on-prem ↔ AWS · already passing traffic
          Week 4
          17 – 23 Aug
          VPN #2 + integration · Azure ↔ AWS IPSec. End-to-end multi-cloud test. Log problems + fixes.
          Week 5
          24 – 30 Aug
          Bonus, report, demo, TEARDOWN · demo to assessors → decommission all by 30 Aug 23:59.

          Team & responsibilities PROVISIONAL

          Working allocation only. The team has not confirmed the report section order or final ownership. These cards show current preparation areas and must not be read as approved section assignments.

          Casper

          ScrumMaster

          Proposed speaking focus: architecture, executive summary, conclusion and CIDR sign-off.

          PROVISIONAL

          Open shared team deck

          Jeff

          Team Member

          Proposed speaking focus: VPC/VNet build, both VPN connections, VGW Ireland and customer gateway Virginia.

          PROVISIONAL

          Open shared team deck

          Bernard

          Team Member

          Proposed speaking focus: delivery, critical path, risk, monitoring, recovery and teardown.

          PROVISIONAL

          Open shared team deck

          Adelene

          Team Member

          Proposed speaking focus: IP addressing, routing, test evidence and the public-subnet web server.

          PROVISIONAL

          Open shared team deck

          Soo Fern

          Team Member

          Proposed speaking focus: Virginia workloads, DNS/DHCP, constraints and lessons learned.

          PROVISIONAL

          Open shared team deck

          Readiness planning · provisional preparation map

          Preparation guide, not ownership sign-off. The team now has one shared presentation deck. Speaking ownership remains provisional until the team confirms the order and Responsibility Assignment Table. Use the shared TEAM REVIEW deck and the working report.

          Casper · proposed focusArchitecture, objectives, CIDR sign-off and project conclusions. Needs team confirmation.
          Jeff · proposed focusNetwork construction, VGW/customer-gateway roles and both VPN implementations. Needs team confirmation.
          Bernard · proposed focusDelivery, security, monitoring, cost, ALB integration, recovery and teardown. Needs team confirmation.
          Adelene · proposed focusIP addressing, subnet design, routing and test evidence. Needs team confirmation.
          Soo Fern · proposed focusVirginia workloads, strongSwan forwarding, constraints and lessons learned. Needs team confirmation.

          Decision record: report order, accountable section owners and demo roles are TBD by the full team, coordinated by Casper, target 29 Jul. Confirmation sequence: agree the report order → assign one accountable owner to each section → update the Responsibility Assignment Table and this page → then add page references and run the readiness check.

          The test to apply after ownership is confirmed. Have someone outside the team open the assigned material at random and ask one question. If the owner must read from the document to answer, more preparation is needed.

          Interactive access is still a single point of failure. The automated refresh uses short-lived OIDC access and the operator scripts are shared, but the demo sessions, SSH keys and recovery path have not been proven from a second device. Backup operator TBD by the team by 12 Aug. Closure requires one end-to-end rehearsal with Bernard observing but not operating.

          4Evidence we are on trackThe artefacts and the proof, measured against the standard above.

          Deliverables · current working artefacts

          Project report

          Group report prepared for team review, 20 rendered pages. It covers all three VPNs, bidirectional Virginia–Azure evidence, the integrated two-AZ ALB recovery design, the fresh 218-second recorded recovery drill, four resource-group-scoped Azure Reader accounts, eight genuine-action chapters, IaC controls and dependency-ordered teardown.

          DOCX 473 KB 20 pp TEAM REVIEW

          Download report

          Technical update complete; team approval pending. 484,174 bytes, updated 30 Jul. School-account deployment is explicitly deferred.

          Group presentation · one shared deck

          One 13-slide team deck replaces the five individual working decks. It covers the three-VPN architecture, delivery, security/monitoring, cost, direct end-to-end tests, the fresh 218-second ALB recovery drill, Azure team access, Q&A and teardown. Every slide contains provisional presenter notes and a source block.

          PPTX 13 slides 68 KB TEAM REVIEW

          Download shared deck

          Current working copy, pending team approval. 69,871 bytes, updated 30 Jul. The speaking map remains explicitly provisional.

          Infrastructure as Code

          The school-account package contains account-neutral CloudFormation YAML for Virginia and Ireland, Azure Bicep for Singapore, VPN #1/#2/#3, two-AZ ALB and recovery, logging/monitoring controls, sanitized strongSwan templates, validation scripts and the staged deployment runbook. Generated tunnel addresses and VPN secrets are supplied at deployment time and never stored in Git.

          Download school IaC package

          Runbook · integrated hub ALB Guide · staged rebuild Runbook · hub recovery Matrix · resource ownership Runbook · school account

          Extracts are reproduced in report sections 5.3.1 to 5.3.3. Full templates are held with the team repository.

          Demonstration recordings · eight chapters

          Eight genuine-action recordings published and verified through CloudFront at 15:29 SGT on 30 July. These are not dashboard walkthroughs: they show authenticated cloud queries, provider-native validation, live private-path tests, a controlled recovery drill and a read-only teardown dry-run as the activities occur. Credentials and VPN secrets are excluded. S3 versioning preserves the replaced objects; every public download matches the approved local hash and byte size.

          1 · Live architecture baselineAuthenticated AWS and Azure identities plus live discovery of the Virginia, Ireland and Singapore cells. ACTION RECORDED
          2 · IaC validationLocal fail-closed controls and provider-native validation of all CloudFormation YAML and Azure Bicep rebuild artefacts; no deployment performed. ACTION RECORDED
          3 · VPN #1 · Virginia to IrelandLive AWS tunnel telemetry, temporary EC2 Instance Connect access, four-packet private ping and HTTP response from the Ireland hub; 0% packet loss, about 69 ms. ACTION RECORDED
          4 · VPN #2 · Singapore to IrelandConnected Azure VPN telemetry plus Azure Run Command ping and HTTP tests to the Ireland hub; 0% packet loss, about 189 ms. ACTION RECORDED
          5 · Direct VPN #3 · Virginia to SingaporeConnected Azure telemetry and bidirectional private-address pings; 0% packet loss in both directions at about 229–230 ms. ACTION RECORDED
          6 · Controlled ALB hub recoveryFresh recorded drill: primary stopped, cold replacement served through DNS and ALB in 218 seconds against the 600-second objective, then primary and zero-capacity steady state restored. PASS · 218 S
          7 · Monitoring, audit and costLive CloudWatch, CloudTrail, both VPC Flow Logs and fail-closed AWS/Azure billing refresh; four alarms OK and no billing query errors. ACTION RECORDED
          8 · Portability and teardown dry-runIaC validation and dependency inventory with the safe borrower-before-owner order. Read-only: no resource was stopped, updated or deleted; school deployment remains deferred. DRY-RUN RECORDED

          Open capture, redaction, S3 and timing plan

          Diagrams · network, recovery, evidence and build sequence

          1 · Core network and VPN topology

          Core designed and rebuildable multi-cloud network topology showing the two AWS VPCs, Azure VNet, nested subnets, strongSwan routed test host, Apache and nginx workloads, two Site-to-Site VPN connections, staged Azure gateway and the blocked transitive routeClick to enlarge

          Designed/rebuildable state, not a live-status claim. The diagram follows the current CloudFormation and Bicep and labels the Azure gateway as staged for the demonstration. The live map and state banner above remain authoritative for what is running now. Both AWS VPCs contain one public subnet; Azure contains one workload subnet and one GatewaySubnet. The strongSwan EC2 owns the Elastic IP and routes bidirectional traffic for the Amazon Linux 2 ICMP/SSH test host. The current Ireland and Singapore workload addresses are labelled as current because they may change after rebuilding. The VGW design remains deliberately non-transitive. Open editable vector · Open Figure 5.0 exactly as submitted in the report.

          2 · Hub-service recovery · accepted 10-minute objective

          Six-stage AWS hub EC2 recovery sequence from detection through cold Auto Scaling, health validation, Route 53 private DNS switching, measurement and restoration, showing a 218-second pass against the 600-second objectiveClick to enlarge

          Resilience boundary and measured result. EventBridge or two failed one-minute EC2 status checks start the cold recovery group, build an Apache replacement in the separate 10.1.2.0/24 recovery subnet, validate EC2 and HTTP health, and move the 30-second private DNS record. The fresh recorded drill passed in 218 seconds against the accepted 600-second objective. This is application-service recovery only: the VPC, VGW and VPN connections remain in place. VPN #2 was live and separately proven before the drill; Azure-client continuity during every recovery transition was not claimed. Open editable vector · Open the dated evidence.

          3 · Live evidence and public-dashboard pipeline

          Live evidence pipeline from authenticated AWS and Azure APIs through freshness-gated GitHub Actions collection, complete and error-aware records, Workers KV and committed fallbacks, Cloudflare Pages Functions and the assessor-facing public dashboardClick to enlarge

          Why the public evidence is trustworthy. GitHub Actions uses AWS and Azure OIDC to collect state normally hourly and the larger inventory approximately every three hours. Complete, error-aware JSON is written to Workers KV as the primary path; committed files are the browser fallback. The Cloudflare watchdog checks KV every five minutes and dispatches a freshness-gated workflow only when needed, while the native GitHub schedule is an independent fallback. The dashboard distinguishes fresh, stale, partial and unreachable: freshness never substitutes for completeness, and no VPN PSK, SSH key or cloud credential is published. Open editable vector.

          4 · Cross-cloud build sequence

          The cross-cloud build order, six steps

          Figure 6.1. The cross-cloud build order. Step 3 is the hinge: AWS only generates the tunnel addresses and pre-shared keys once the VPN connection exists, so neither cloud can be completed first. The seam has to be walked once, in order, by hand.

          5 · Supporting service catalogue · report Figure 5.7

          The VPN environment and the services around it: infrastructure as code above, the two Site-to-Site VPN connections in the middle, and the detect, record and cost-control services belowClick to enlarge

          Figure 5.7. The VPN environment and the services around it. The two VPN connections are the assets being protected; AWS supplies two tunnel endpoints for each connection. The current templates define both VPC Flow Logs, the multi-region CloudTrail, SNS and both VPN alarms. AWS Budgets and the CloudWatch dashboard remain separate account-level controls. The retained central S3 bucket is intentionally left after stack deletion until evidence is preserved and it is explicitly emptied and removed. Current deployed ownership must be proved from describe-stack-resources; until a complete inventory succeeds, it is TBD rather than assumed. Open the ownership matrix.

          Evidence · dated console captures taken while the system was live

          These screenshots prove the tested topology at the time they were captured. Some identifiers changed during later rebuilds; the live map and current-state banner above are authoritative for the current deployment.

          Both VPN connections available on the same VGW

          Both VPN connections Available on the same Virtual Private Gateway. This single frame is the proof of the hub topology.

          VPN 1 tunnel up

          VPN #1 to the simulated on-premises cell. Customer gateway 52.201.145.0, tunnel 34.247.143.4 Up.

          VPN 2 tunnel up

          VPN #2 to Azure. Customer gateway 40.119.233.66, tunnel 52.49.122.112 Up. This is the cross-cloud link.

          Azure route propagated to VGW

          The Azure network 10.2.0.0/16 propagated to the VGW, Active, Propagated: Yes. Learned automatically, not typed in.

          AWS route table with four routes

          The hub route table. Four routes, with both remote networks learned through Virtual Private Gateway propagation.

          Test results · six tests, five pass, one expected fail

          1. VPN #1 data planePASS · on-prem to 10.1.1.61, 0% loss, ttl 254, ~69 ms
          2. VPN #1 IPSec SAPASS · ESTABLISHED, INSTALLED, ESP in UDP (NAT traversal confirmed)
          3. VPN #2 data planePASS · Azure 10.2.1.4 to 10.1.1.61, 0% loss, ttl 254, ~189 ms
          4. VPN #2 application layerPASS · HTTP from Azure returned the AWS web server page
          5. Azure web serverPASS · reachable on its public IP
          6. Transitive routingEXPECTED FAIL · on-prem cannot reach Azure, 100% loss, correct by design

          Test 4 is the strongest single result in the project. A real application request crossed an encrypted tunnel between two different cloud providers and returned a response. Test 6 was run deliberately: knowing where an architecture stops is part of knowing the architecture.

          5ControlsWhat stops this going wrong quietly.

          Security & monitoring · detective controls live (Bernard · Security/Monitoring)

          CloudWatch alarms
          4 OK
          Hub recovery RTO, hub EC2 health, VPN #1 and VPN #2 all returned to OK after the recorded drill.
          CloudTrail
          logging
          multi-region · log-file validation on
          VPC Flow Logs
          2 × ACTIVE
          both AWS VPCs → S3, objects delivering
          SNS alerts
          confirmed
          email subscription active on the topic
          EG334S-Hub-Recovery-Over-10-Minutes
          OKHubRecoverySeconds threshold: 600 seconds · fresh recorded drill: 218 seconds
          EG334S-Hub-StatusCheckFailed
          OKPrimary hub EC2 system and instance health checks are normal.
          EG334S-VPN1-TunnelDown
          OKvpn-065016cdb8d06a873 · observed tunnel 54.77.211.189 · metric math covers both AWS tunnel endpoints
          EG334S-VPN2-TunnelDown
          OKvpn-03dcb879c286f5ecd · metric math covers both AWS tunnel endpoints
          VPN #2 was live for the 30 Jul action recording and the alarm returned to OK. The live banner remains authoritative if the environment changes after this dated evidence.
          CloudTrail
          EG334S-audit-trail · IsLogging: true · log-file validation enabled
          VPC Flow Logs
          fl-0963dc40ce2f1c698 (eu-west-1) + fl-0ef531e2d2d6d1173 (us-east-1)
          Destination: s3://eg334s-flowlogs-980195619820
          CloudWatch dashboard
          EG334S-MultiCloud-Monitoring (eu-west-1)
          Tunnel state · tunnel traffic · EC2 CPU
          SSH posture
          The 25 Jul deployment is restricted to one administrative /32 on the on-premises hosts, with no public SSH ingress on the hub server. The published CloudFormation templates have no permissive default: deployment requires an explicit valid /32, and scripts/validate-iac.sh rejects an open administrator default.

          Security groups decide what is allowed. Alarms, trails and flow logs decide what is noticed. The VPN #2 alarm firing on a real outage is stronger proof than a pair of green ticks: it shows the detection path works end to end, from metric to alarm to SNS to inbox.

          AWS hub automated recovery drill · PASS

          Decision: protect the AWS hub application without duplicating or replacing the VPN infrastructure. The accepted objective is to restore the hub web service within 10 minutes of a detected EC2 failure. This is a controlled application-service recoverability drill—not full High Availability, regional disaster recovery, or VGW/VPN failover.

          PASS · 3m 38s RTO ≤ 10 min standby 0 EC2 NOT FULL HA Azure proven separately

          Drill objective and measured result
          Recovery objective
          ≤ 10 min
          Accepted RTO for loss of the hub EC2 application
          End-to-end test
          3m 38s
          218 seconds from stop request to confirmed HTTP 200
          Automation interval
          187.95 s
          CloudWatch metric: controller start to private DNS publication
          Normal standby
          0 EC2
          Cold Auto Scaling group · maximum one replacement
          Recovery design and cost choices
          DESIGN CHOICE

          Replace the service, preserve the network

          The failed component is the hub web-server EC2 instance. The recovery stack therefore imports the existing VPC, security group and propagated route table, then launches a replacement in eu-west-1b. It does not create or modify the Virtual Private Gateway, customer gateways or either Site-to-Site VPN connection.

          This keeps the blast radius small. The working VPN control plane remains in place while only the stateless application host is replaced.

          COST CHOICE

          Cold standby, automated launch

          The recovery Auto Scaling group normally has desired capacity zero. Approximate steady-state incremental cost is US$0.70/month: one Route 53 private hosted zone and two standard CloudWatch alarms. Lambda, EventBridge, Parameter Store and DNS-query usage are negligible at this project scale.

          A replacement t3.micro, its 8 GiB gp3 disk and public IPv4 are billed only while recovery capacity is running. The stopped primary continues to incur its EBS storage cost.

          Seven-step automated recovery sequence

          What happens when the primary hub fails

          1 · DetectEventBridge reacts to the primary entering stopped or terminated. A separate two-of-two, one-minute StatusCheckFailed alarm covers EC2 system or instance failure.
          2 · StartLambda records the recovery start time and changes eg334s-hub-recovery-HubRecoveryAsg from desired capacity 0 to 1.
          3 · RebuildThe launch template creates a hardened Amazon Linux replacement, installs Apache, enables automatic service restart and publishes the project health page.
          4 · GateTraffic is not moved merely because the instance says running. The DNS updater waits until both EC2 system and instance status checks report ok.
          5 · PublishThe 30-second private record hub.eg334s.internal changes to the replacement private IP. A VPC-attached Lambda then requests the service through that stable name and requires HTTP 200.
          6 · MeasureEG334S/Recovery:HubRecoverySeconds records the automated interval. EG334S-Hub-Recovery-Over-10-Minutes alarms if the measured recovery exceeds 600 seconds.
          7 · RestoreAfter the original primary is healthy, DNS returns to 10.1.1.61 and recovery capacity returns to zero. The order matters: never remove the replacement before primary health and DNS restoration are confirmed.
          Dated evidence and claim boundary
          AWS HUB AUTOMATED RECOVERY DRILL · 30 JUL 2026

          Primary stop initiated: i-074bf0cb910a99e02 at 14:48:00 SGT. Replacement: i-0007949b2f976f068 at 10.1.2.154. Result: private DNS and the public ALB returned HTTP 200 after 218 seconds. The primary was restarted, private DNS returned to 10.1.1.61, recovery capacity returned to zero, and all four project alarms were OK.

          Claim boundary. This proves recovery of the stateless hub EC2 application inside the existing VPC. It does not prove recovery from loss of eu-west-1, the VPC, the VGW, Route 53 or stateful application data. It also does not preserve the instance's public IPv4 address. Team tests should use hub.eg334s.internal from associated private networks rather than treating a public IP as the failover endpoint.

          Azure scope. VPN #2 and direct VPN #3 were live and proven with genuine private-path tests before the recovery drill. The 218-second metric measures AWS hub application recovery through private DNS and the public ALB; it does not claim uninterrupted Azure-client service during every transition.

          Dated test evidence Recovery runbook

          Cost monitoring · guardrails active (Bernard · Security/Monitoring)

          Checking billing APIs…

          Azure · month to date
          loading
          Cost Management actual cost · refreshed with the live state feed
          AWS · month to date
          loading
          Net billed and usage before credits · refreshed with the live state feed
          Idle now (torn down)
          ~$0.44/day
          Azure = 2 static IPs + disk · ~$7.85/day only while rebuilt for the demo
          Guardrails
          3 budgets
          AWS $40/mo + $3/day · Azure $25/mo
          AWS Budgetsactive · EG334S-Monthly-40 + EG334S-Daily-Tripwire-3 → email
          Cost Anomaly Detectionpresent but incapable of firing at this scale. It is the AWS Default-Services-Monitor, not one configured for this project, and its subscription requires an anomaly of ≥ $100 AND ≥ 40%. Peak spend here is about $7.85/day, so the $100 condition cannot be met. Its notification address also differs from the SNS topic address; confirm which inbox is actively monitored.
          Daily still-running checkscheduled · 09:00 SGT, push + email, teardown commands attached
          Auto-stop on breachnot implemented, but it is available. bernard-admin holds AdministratorAccess and simulate-principal-policy returns allowed for iam:CreateRole, iam:PassRole, budgets:ModifyBudget and ec2:StopInstances. An earlier note claimed lab IAM blocked Budget Actions. That was true of the previous account and of the agent guardrail, not of this account.
          Azure budgetactive · EG334S-Azure-Guardrail-25, $25/mo, alerts 40 / 75 / 100% + forecast → email
          Azure spending limitOff, and it cannot be turned on. The spending limit is not a switch we left off. Azure only offers it on credit-based subscriptions (free trial, Azure for Students, Visual Studio), where its job is to disable the subscription once the included credit runs out rather than start charging a card. This subscription is PayAsYouGo_2014-09-01, which has no included credit to protect, so Azure reports the limit as Off and offers no way to enable it. That is why spend here is uncapped and why a budget with alerts is the only native control available. For a hard cap you would need a budget wired to an action group that triggers an automation runbook to deallocate resources, which is the Azure equivalent of AWS Budget Actions.
          Cost-allocation tagactive · the Project tag was activated on 25 Jul and now reports Status: Active. An earlier note claimed it was not activatable from a linked account. That was wrong, and it was never tested. Note the lag: activation only applies to usage recorded from that point on, so it will not retrospectively split earlier spend.

          How these numbers were obtained. Actual month-to-date spend comes from the Azure Cost Management API and AWS Cost Explorer. Projected daily run rates are estimates calculated from the Azure Retail Prices API and AWS Price List API against the real resource inventory. At the 25 Jul rate check, Azure VPN Gateway VpnGw1AZ was $0.21/hr, the Azure VM Standard_D2als_v7 was $0.101/hr, and each AWS Site-to-Site VPN connection was $0.05/hr.

          Actual versus projected, and why they differ so much. The earlier $283 figure was a projection of what leaving everything running to the deadline would have cost. The current month-to-date values are displayed above from the billing APIs. Azure remains lower than the projection because the gateway runs only in short bursts—built, verified and torn down again—rather than continuously.

          How fresh can this be? Azure Cost Management returns same-day data, though the current day keeps trickling in for several hours. AWS Cost Explorer lags a day or more and flags its figures as Estimated. So "live cost" does not exist on either cloud. What is real time is the resource inventory, which is what actually matters for catching something left running. Any team member with the documented read-only access can run ./scripts/refresh-cost.sh; it fails closed if either cloud login or a required billing field is unavailable.

          Control lesson. The first guardrails covered AWS even though credits reduce its net bill to approximately zero, while the Pay-As-You-Go Azure subscription carried the cash cost. Azure now has its own monthly budget and alerts.

          Guardrails are alert-only: the cloud emails you, you act. The cheapest guardrail is still teardown · delete-stack on AWS and az group delete on Azure between work sessions.

          Risk watch · the things that actually bite

          Owner means whoever can actually act. The programme manager tracks every risk on this list and chases all of them, but only owns the ones where he holds the lever. A risk assigned to someone who cannot discharge it is not managed, it is just parked.

          !
          Key-person dependency · OPEN · automation uses short-lived OIDC access and the verification tools are now shared in this repository, but the interactive demo sessions, SSH keys and recovery access have not been proven from a second device. Backup operator: TBD by the team by 12 Aug. Closure evidence is a complete rehearsal in which Bernard does not operate the environment.
          !
          Section ownership · DECISION PENDING · the current report order and owner mapping are a working proposal. This is deliberately labelled provisional rather than treated as fact. Decision owner: full team; coordinator: Casper; target: 29 Jul. Closure requires one accountable owner per section, an approved Responsibility Assignment Table and matching dashboard/readiness records.
          !
          Shared group presentation · TEAM REVIEW READY · one 13-slide deck now carries the common narrative, provisional speaking map, assessor Q&A and sourced speaker notes. Next action: each member validates the technical statements they may present; the full team confirms the order and ownership before rehearsal.
          !
          Live cross-cloud cost exposure · VPN #2 and direct VPN #3 were live during the 30 Jul action recordings. Action: keep the environment only for the approved demonstration window, monitor the live billing cards, and use the dependency-ordered teardown after assessment. Owner: Bernard.
          !
          Decommission slippage · the brief states a real mark penalty for AWS resources left after 30 Aug 23:59 SGT. The templates now define Flow Logs, CloudTrail, SNS and VPN alarms, while the central evidence bucket is deliberately retained and budgets/dashboard remain manual. Current deployed ownership is TBD until a complete inventory and stack-resource reconciliation succeed. The Cloudflare evidence-site retention decision is also TBD by the team/lecturer. Owner: Bernard; second verifier: TBD. See the ownership matrix.
          !
          Public working artefacts · DECISION PENDING · the team-review report and shared group deck are publicly downloadable under the user's explicit approval. The full team should still confirm institutional rules and final submission status; both artefacts remain clearly labelled non-final.
          Closed · AWS hub automated recovery drill · the accepted 10-minute objective was re-exercised and recorded on 30 Jul. A controlled primary stop launched a cold replacement in a second Availability Zone; stable private DNS and the public ALB returned HTTP 200 in 218 seconds (3m 38s). Primary health, DNS and zero standby capacity were restored, and all four project alarms were OK. This proves EC2 application recoverability, not full HA, VGW/VPN failover or regional DR. Open the dated evidence.
          Closed · report administration · four supplied admin numbers are present; Bernard's remains explicitly TBC. “Adelene” is corrected, pagination defects are removed, and the current team-review copy renders to 20 pages.
          Closed · overlapping CIDRs (locked at M1), evidence debt (captured while live), gold-plating (core proven before monitoring was added), cross-region teardown (documented per region and per subscription).

          Claim verification · dated evidence record

          This is a historical evidence record, not the hourly operational reading. Current tunnel, alarm, compute and cost summaries come from the live banner and map. Infrastructure facts below were captured after the 25 Jul rebuild; working-artefact metadata was updated on 28 Jul.

          AWS identifiersre-verified 25 Jul after rebuild · vgw-0bf466d711cc93ac4, VPN #1 vpn-065016cdb8d06a873, VPN #2 vpn-03dcb879c286f5ecd, both VPN connections, both flow logs. Verified with describe-vpn-connections, describe-customer-gateways, describe-flow-logs.
          Addressesconfirmed 25 Jul · web server 10.1.1.61, VPN #1 tunnel 54.77.211.189, on-prem Elastic IP 3.214.123.92, VPN #2 tunnel 34.243.188.162, Azure gateway IP 40.119.233.66.
          Design assertionsconfirmed 25 Jul · SourceDestCheck=false on the strongSwan instance, GatewaySubnet 10.2.255.0/27 with nsg: None, workload subnet 10.2.1.0/24 with its NSG attached, local network gateway configured for the current AWS VPN #2 endpoint and 10.1.0.0/16.
          Monitoring and budgetsconfirmed 28 Jul · all four alarms exist (hub recovery duration, hub EC2 health, VPN #1 and VPN #2), CloudTrail IsLogging: true, SNS has a confirmed email subscription, both AWS budgets remain at $40 and $3, and the Azure budget remains at $25.
          Cost arithmeticrate basis recomputed 25 Jul · $0.21/hr + $0.101/hr + 2 IPs + disk = approximately $7.85/day Azure when fully rebuilt for the demo. Torn down it idles at approximately $0.44/day. Current month-to-date billing values are shown in the live cost cards above.
          Latencyre-measured 25 Jul · VPN #1 returned 0% loss at 70.3 ms average with ttl=254, against the ~69 ms recorded earlier. Within normal jitter, same path.
          Working artefactspending team approval · report: 484,174 bytes and 20 rendered pages. One 13-slide shared team deck with provisional presenter notes and source blocks: 69,871 bytes.
          AWS control auditrecorded verification · 30 Jul · expected EC2 instances, active two-AZ ALB, recovery ASG restored to desired 0, VPN #1 and VPN #2 live during the action window, all four project alarms OK, CloudTrail logging and both VPC Flow Logs active. Open dated audit.
          Live state, routes, ownership and costPASS · 30 Jul, 17:42–17:47 SGT · authenticated read-only capture confirmed both AWS VPNs, both Azure connections, three running AWS instances, the running Azure VM, healthy two-AZ ALB, four alarms OK, effective routes/security controls, all active AWS stack and Azure deployment owners, and a pre-parking billing baseline. No infrastructure changed. Open consolidated evidence.
          Automated recovery drillPASS · 30 Jul · controlled stop of i-074bf0cb910a99e02; replacement i-0007949b2f976f068 served HTTP through private DNS and the public ALB at 10.1.2.154 in 218 seconds. Primary, DNS and zero-capacity steady state were restored; all four alarms were OK. This is recoverability evidence, not full HA. Evidence and claim boundary.

          Assessment documents verified 28 Jul 2026. The assessment weightings, scope, report requirements and both rubrics were checked against the supplied official PDFs. The report content order shown in the brief is an example for reference, not a mandatory sequence. The current working report was rendered and checked at 20 pages; the shared deck was rendered across all 13 slides. Both remain subject to team review. One infrastructure limitation remains unverified: the orphaned IAM role sits on the previous lab account, which this session has no access to, so its continued existence is reported, not observed.

          Re-run this before the demo. Infrastructure claims decay. The shared ./scripts/verify-teardown.sh and ./scripts/refresh-cost.sh tools are now in this repository. Both refuse to report success if authentication or a required query fails. Use the provisional evidence register to replace historical captures after the 25 Aug rebuild.

          6How it endsDecommissioning. This is where marks are lost.

          Teardown checklist · run in every region before 30 Aug 23:59

          On-prem sim AWS us-east-1

          • EC2 private server
          • Customer gateway + VPN connection
          • VPC, subnets, route tables, IGW
          • Elastic IPs released
          • Security groups + key pairs
          • CloudFormation stack deleted

          AWS public AWS eu-west-1

          • Delete eg334s-hub-recovery first
          • EC2 web server
          • Virtual private gateway + VPN conn
          • CloudWatch / CloudTrail / Flow Logs
          • VPC, subnets, route tables, IGW
          • Elastic IPs released
          • CloudFormation stack deleted

          Azure Southeast Asia

          • az group delete --name eg334s-rg --yes
          • Removes VPN gateway (the big cost)
          • Removes connection + local network gateway
          • Removes VNet, NSG, VM, disks, public IPs
          • az resource list --tag Project=EG334S returns empty

          AWS order matters: delete eg334s-vpn2-aws before eg334s-vpn1-awspublic, it borrows that VGW.

          Note: an orphaned IAM role from the locked-down lab account (eg334s-vpn1-onprem-SsmRole-*) can't be self-deleted · flag to lab admin. IAM is global, not regional.

          Two controls that protect the teardown

          1 · DEPENDENCY ORDER

          Delete the borrower before the owner

          eg334s-vpn2-aws borrows the virtual private gateway owned by eg334s-vpn1-awspublic. Delete VPN #2 first so the owner stack can remove the gateway cleanly.

          2 · RETAIN EVIDENCE

          Prove that nothing remains

          After teardown, run the authenticated verification script across both AWS regions and Azure. Retain its PASS / FAIL / ERROR transcript as the M6 completion evidence.

          Technical note

          The historical deployment used a separate VPN #2 stack that borrowed the hub VGW by parameter, so CloudFormation did not enforce deletion order. The current published hub template can stage VPN #2 inside the same stack. Which model is live is TBD until stack-resource inventory succeeds; follow the observed model, not an assumption.

          The shared verification script checks the project-named AWS resources returned by the complete inventory plus every resource in eg334s-rg, and fails closed with ERROR if authentication or any required query fails. Account-level budgets, the CloudWatch dashboard, key pairs, legacy IAM and the Cloudflare evidence site remain explicit manual/TBD controls in the ownership matrix.

          Shared evidence record · technical milestones are source-controlled, not browser-local. Times in Asia/Singapore. Built for EG334S. Live status is in the banner and map above; the update stamp is below.

          Ambience Mozart · 5 recordings
          Track 1 of 5 · in order, loops
          Checking playback mode

          Le nozze di Figaro, K. 492, Act III, No. 21, Sull’aria … Che soave zeffiretto, the Letter Duet. Five interpretations. Press play above. Full tracks need an active Spotify session in this browser, otherwise Spotify serves previews. Turn this off before the assessed demo.