On-prem VPC ↔ AWS public VPC ↔ Azure VNet, joined by two Site-to-Site VPN connections
All dates are SGT. Sessions 1 and 2 are Sync-Physical; sessions 3 to 15 are Sync-Online. Session 8 on 11 Aug is the National Day public-holiday make-up lesson. The VPN #2 rebuild on 25 Aug is a separate internal action so the 26 Aug assessment can run on live infrastructure. Teardown must be verified in us-east-1, eu-west-1, Azure and Cloudflare.
60% is individual (Checkpoint + Presentation). Build the system as a team; own and defend your own section alone.
| Performance indicator | Excellent (top band) | What it takes |
|---|---|---|
| Organisation & Content (60) Organisation & supporting materials; Content |
Agenda exists, coherent & interesting sequence; supporting materials used innovatively & explained in context; can explain all details, constraints and work-arounds. | Have a clear agenda; every diagram/screenshot explained in context; be able to explain the whole project's details, limits and how you worked around them. |
| Presentation Skills (40) Delivery; Q&A |
Interesting, eloquent, enthusiastic delivery (not heavily scripted); handles all Q&A well and anticipates questions. | Rehearse so you're not reading slides; pre-empt likely assessor questions and prepare answers. |
| Performance indicator | Excellent (top band) | What it takes |
|---|---|---|
| Report Presentation (20) Writing; Presentation & supporting materials |
Exceptionally clear, precise, concise English; few typos; professional layout; all illustrations well formatted. | Proofread hard; consistent styles/margins; clean, labelled figures and screenshots. |
| Technical Content (30) Organisation & structure; Literature survey; Quality of analysis |
Structure entirely correct, all sections placed; exemplary range of references; well-informed, authoritative discussion of a complex problem with depth. | Use a logical, complete structure; cite an exemplary range of authoritative references; show reasoned technical analysis, not just description. |
Verified on 27 Jul 2026 against the supplied official Presentation Assessment Rubrics and Report Assessment Rubrics PDFs. Presentation is marked out of 100 (60 + 40); the report rubric has 50 rubric marks (20 + 30) and the report contributes 20% of the module assessment.
Not yet confirmed by the team. The brief labels its contents page as a sample for reference only, so the order and owner chips shown here are planning placeholders. Use the matrix to develop the proposal; do not treat it as the approved Responsibility Assignment Table. Changes save to the shared planning matrix that every visitor sees.
10.0.0.0/16 present in the eu-west-1 route table, learned by propagation, not typed40.119.233.66 kept) + connection the day before the demo (see REBUILD.md phase D)curl http://10.1.1.61 from the Azure VM returned the AWS web server page, now served by the template itself · a real workload across clouds. VPN #1 was proven with ICMP, not HTTP, so this row does not apply to it.10.0.0.0/16 from VPN #1 by propagation; 10.2.0.0/16 re-propagates automatically when VPN #2 is rebuilt.Known by design: on-prem cannot reach Azure directly (tested: 100% loss). An AWS VGW does not do transitive routing between two VPN connections. Topology is hub-and-spoke with AWS public as the hub, matching the brief. Full mesh would need a Transit Gateway.
AWS public cloud is the hub: its single VGW terminates both VPN connections, each with two AWS-provided tunnel endpoints. Latencies are measured round-trip times from the ICMP tests, and both replies returned ttl=254, which proves the packets crossed a gateway rather than being answered locally. The difference is purely physical distance: trans-Atlantic versus half-way round the world. Azure was placed in Singapore rather than Ireland so the assessed demo runs at local latency, the trade being a longer VPN #2 connection that is invisible to the viewer. Note that the VGW does not perform transitive routing, so on-prem cannot reach Azure directly. That is hub-and-spoke by design. Teardown must clear us-east-1, eu-west-1 and Azure separately. Current state: VPN #2 is torn down to control cost. It was rebuilt and verified end-to-end on 26 Jul, then the Azure gateway, the connection and the VM were removed, leaving Azure at about $0.44/day. Both static IPs are kept, including 40.119.233.66 that the AWS customer gateway is pinned to, so rebuilding for the demo is a ~30-minute Azure-only job (REBUILD.md phase D).
Nothing on this list is typed by hand. The
live-data workflow runs scripts/inventory.sh, which only ever calls describe
and list, and refreshes this feed automatically. Every required query records success or
failure; a failed table is shown as unknown, never none. Anything here that is not inside a
CloudFormation stack or the Azure resource group has to be deleted by hand at teardown,
because delete-stack and az group delete will not touch it.
Green = done. Amber = AWS side done, Azure pending. Grey = not started. Anything off this chain (bonus, report polish, slides) is parallel work you sacrifice first if time runs short.
This is a technical milestone count, not an overall completion percentage. Status is defined in the shared page source and changes only with reviewed evidence.
Working allocation only. The team has not confirmed the report section order or final ownership. These cards show current preparation areas and must not be read as approved section assignments.
Current preparation focus: architecture, executive summary, conclusion and CIDR sign-off.
PROPOSED FOCUSCurrent preparation focus: VPC/VNet build, both VPN connections, VGW Ireland and customer gateway Virginia.
PROPOSED FOCUSCurrent preparation focus: delivery, critical path, risk, teardown, CloudWatch, CloudTrail and VPC Flow Logs.
PROPOSED FOCUSCurrent preparation focus: IP addressing, routing and the public-subnet web server.
PROPOSED FOCUSCurrent preparation focus: Virginia private server, DNS/DHCP and the problems-and-solutions log.
PROPOSED FOCUSPreparation guide, not ownership sign-off. Exact sections and page numbers are deliberately omitted until the team confirms the report order and Responsibility Assignment Table. Select a member's proposed focus to download the current team-review report.
SourceDestCheck and route propagation. Needs team confirmation.Decision record: report order, accountable section owners and demo roles are TBD by the full team, coordinated by Casper, target 29 Jul. Confirmation sequence: agree the report order → assign one accountable owner to each section → update the Responsibility Assignment Table and this page → then add page references and run the readiness check.
The test to apply after ownership is confirmed. Have someone outside the team open the assigned material at random and ask one question. If the owner must read from the document to answer, more preparation is needed.
Interactive access is still a single point of failure. The automated refresh uses short-lived OIDC access and the operator scripts are shared, but the demo sessions, SSH keys and recovery path have not been proven from a second device. Backup operator TBD by the team by 12 Aug. Closure requires one end-to-end rehearsal with Bernard observing but not operating.
Group report prepared for team review, 21 rendered pages. Cover, fixed contents, a provisional responsibility assignment table, sections 1 to 8 and Annex A bibliography. Includes the design basis and literature review, both drawn diagrams, five console screenshots, infrastructure-as-code extracts, the strongSwan configuration, all six tests, the AWS hub automated recovery drill and the complete teardown checklist.
DOCX 274 KB 21 pp TEAM REVIEW
This is the current working copy, pending team review and approval. Public distribution is TBD by the team and institutional rules, decision due 29 Jul. MD5 81287885425387e67c9caba089f2944a, 280,087 bytes, 28 Jul. If the copy you are holding does not match, it is an older draft. Check with md5 <file> on a Mac or certutil -hashfile <file> MD5 on Windows.
Draft markings and the red bibliography note have been removed, page numbers and contents are fixed for this copy, Test 5 is explicitly documented, and the technical claims and sources have been corrected. Team feedback may still change the submission.
Bernard's individual deck prepared for team review, 13 slides, with presentation cues and source lists in the speaker notes on every slide. Covers delivery and critical path, security and monitoring, cost guardrails, constraints and workarounds, the transitive-routing finding, the AWS hub automated recovery drill, test evidence and anticipated Q&A.
PPTX 13 slides 69 KB TEAM REVIEW
This is the current working copy, pending team review and approval. Public distribution is TBD by the team and institutional rules, decision due 29 Jul. MD5 37b6a7f368611006dca1bcdc2ac0a462, 70,324 bytes, 28 Jul. The current version includes the AWS hub automated recovery drill and states explicitly that it is not full HA.
The deck distinguishes two VPN connections from four AWS tunnel endpoints, treats TTL as supporting evidence and states AWS cost after credits. Its resource-ownership wording must be reconciled against the final complete inventory before submission.
All three cells are declared as code. The two CloudFormation templates and Azure Bicep now cover the core network and compute, both VPN connections, both VPC Flow Logs, CloudTrail, SNS and the VPN alarms. The generated tunnel addresses and VPN secret are supplied through staged parameters rather than stored in Git. Budgets, the dashboard feed and its scheduler remain separate operational controls.
Guide · staged rebuild Runbook · hub recovery Matrix · resource ownership
Extracts are reproduced in report sections 5.3.1 to 5.3.3. Full templates are held with the team repository.
Click to enlarge
Professional architecture view. Projector-ready topology with the assessed service names, CIDRs, gateway relationships, VPN types, routing modes, measured latency and monitoring controls. Each cell remains an isolated private network; AWS is the hub because its Virtual Private Gateway terminates both Site-to-Site VPN connections. The red path records the tested design boundary: the VGW does not provide transitive routing between VPN #1 and VPN #2. This is a dashboard presentation view; open Figure 5.0 exactly as submitted in the report.
Figure 6.1. The cross-cloud build order. Step 3 is the hinge: AWS only generates the tunnel addresses and pre-shared keys once the VPN connection exists, so neither cloud can be completed first. The seam has to be walked once, in order, by hand.
Click to enlarge
Figure 5.7. The VPN environment and the services around it. The two VPN connections are the assets being protected; AWS supplies two tunnel endpoints for each connection. The current templates define both VPC Flow Logs, the multi-region CloudTrail, SNS and both VPN alarms. AWS Budgets and the CloudWatch dashboard remain separate account-level controls. The retained central S3 bucket is intentionally left after stack deletion until evidence is preserved and it is explicitly emptied and removed. Current deployed ownership must be proved from describe-stack-resources; until a complete inventory succeeds, it is TBD rather than assumed. Open the ownership matrix.
These screenshots prove the tested topology at the time they were captured. Some identifiers changed during later rebuilds; the live map and current-state banner above are authoritative for the current deployment.
Both VPN connections Available on the same Virtual Private Gateway. This single frame is the proof of the hub topology.
VPN #1 to the simulated on-premises cell. Customer gateway 52.201.145.0, tunnel 34.247.143.4 Up.
VPN #2 to Azure. Customer gateway 40.119.233.66, tunnel 52.49.122.112 Up. This is the cross-cloud link.
The Azure network 10.2.0.0/16 propagated to the VGW, Active, Propagated: Yes. Learned automatically, not typed in.
The hub route table. Four routes, with both remote networks learned through Virtual Private Gateway propagation.
Test 4 is the strongest single result in the project. A real application request crossed an encrypted tunnel between two different cloud providers and returned a response. Test 6 was run deliberately: knowing where an architecture stops is part of knowing the architecture.
TunnelState < 1.0. This is the monitoring proving itself. The alarm detected the real Azure-side teardown and the SNS email fired. It should return to OK once rebuild-azure-vpn.sh restores the Azure gateway and connection; current aggregate alarm counts are shown in the live banner.IsLogging: true · tamper-evident via log-file validations3://eg334s-flowlogs-980195619820/32 on the on-premises hosts, with no public SSH ingress on the hub server. The published CloudFormation templates now have no permissive default: deployment requires an explicit valid /32, and scripts/validate-iac.sh rejects an open administrator default.Security groups decide what is allowed. Alarms, trails and flow logs decide what is noticed. The VPN #2 alarm firing on a real outage is stronger proof than a pair of green ticks: it shows the detection path works end to end, from metric to alarm to SNS to inbox.
Decision: protect the AWS hub application without duplicating or replacing the VPN infrastructure. The accepted objective is to restore the hub web service within 10 minutes of a detected EC2 failure. This is a controlled application-service recoverability drill—not full High Availability, regional disaster recovery, or VGW/VPN failover.
PASS · 4m 04s RTO ≤ 10 min standby 0 EC2 NOT FULL HA Azure path not exercised
The failed component is the hub web-server EC2 instance. The recovery stack therefore imports the existing VPC, security group and propagated route table, then launches a replacement in eu-west-1b. It does not create or modify the Virtual Private Gateway, customer gateways or either Site-to-Site VPN connection.
This keeps the blast radius small. The working VPN control plane remains in place while only the stateless application host is replaced.
The recovery Auto Scaling group normally has desired capacity zero. Approximate steady-state incremental cost is US$0.70/month: one Route 53 private hosted zone and two standard CloudWatch alarms. Lambda, EventBridge, Parameter Store and DNS-query usage are negligible at this project scale.
A replacement t3.micro, its 8 GiB gp3 disk and public IPv4 are billed only while recovery capacity is running. The stopped primary continues to incur its EBS storage cost.
stopped or terminated. A separate two-of-two, one-minute StatusCheckFailed alarm covers EC2 system or instance failure.eg334s-hub-recovery-HubRecoveryAsg from desired capacity 0 to 1.running. The DNS updater waits until both EC2 system and instance status checks report ok.hub.eg334s.internal changes to the replacement private IP. A VPC-attached Lambda then requests the service through that stable name and requires HTTP 200.EG334S/Recovery:HubRecoverySeconds records the automated interval. EG334S-Hub-Recovery-Over-10-Minutes alarms if the measured recovery exceeds 600 seconds.10.1.1.61 and recovery capacity returns to zero. The order matters: never remove the replacement before primary health and DNS restoration are confirmed.Primary stopped: i-074bf0cb910a99e02 at 15:26:40 SGT. Replacement: i-03d63eb39288c7419 at 10.1.2.225. Result: hub.eg334s.internal returned HTTP 200 after 244 seconds. The primary was restarted, private DNS returned to 10.1.1.61, recovery capacity returned to zero, and both VPN connections remained available.
Claim boundary. This proves recovery of the stateless hub EC2 application inside the existing VPC. It does not prove recovery from loss of eu-west-1, the VPC, the VGW, Route 53 or stateful application data. It also does not preserve the instance's public IPv4 address. Team tests should use hub.eg334s.internal from associated private networks rather than treating a public IP as the failover endpoint.
Why Azure was not exercised in this recovery test. VPN #2 is intentionally parked between demonstrations to control cost. An Azure end-to-end client test would require rebuilding the VpnGw1AZ gateway and connection, waiting roughly 30 minutes for provisioning, and temporarily returning the environment to its approximately US$7.85/day fully rebuilt run rate instead of the approximately US$0.44/day parked rate. That expense was not necessary to test the accepted failure boundary—the AWS hub EC2 application—because service recovery could be verified from inside the associated private AWS network. Therefore the Azure-to-recovered-hub data path is correctly recorded as not exercised during this test, not as passed. It should be included in the planned full cross-cloud rehearsal after VPN #2 is rebuilt for the assessment.
Checking billing APIs…
Default-Services-Monitor, not one configured for this project, and its subscription requires an anomaly of ≥ $100 AND ≥ 40%. Peak spend here is about $7.85/day, so the $100 condition cannot be met. Its notification address also differs from the SNS topic address; confirm which inbox is actively monitored.bernard-admin holds AdministratorAccess and simulate-principal-policy returns allowed for iam:CreateRole, iam:PassRole, budgets:ModifyBudget and ec2:StopInstances. An earlier note claimed lab IAM blocked Budget Actions. That was true of the previous account and of the agent guardrail, not of this account.PayAsYouGo_2014-09-01, which has no included credit to protect, so Azure reports the limit as Off and offers no way to enable it. That is why spend here is uncapped and why a budget with alerts is the only native control available. For a hard cap you would need a budget wired to an action group that triggers an automation runbook to deallocate resources, which is the Azure equivalent of AWS Budget Actions.Project tag was activated on 25 Jul and now reports Status: Active. An earlier note claimed it was not activatable from a linked account. That was wrong, and it was never tested. Note the lag: activation only applies to usage recorded from that point on, so it will not retrospectively split earlier spend.How these numbers were obtained. Actual month-to-date spend comes from the Azure Cost Management API and AWS Cost Explorer. Projected daily run rates are estimates calculated from the Azure Retail Prices API and AWS Price List API against the real resource inventory. At the 25 Jul rate check, Azure VPN Gateway VpnGw1AZ was $0.21/hr, the Azure VM Standard_D2als_v7 was $0.101/hr, and each AWS Site-to-Site VPN connection was $0.05/hr.
Actual versus projected, and why they differ so much. The earlier $283 figure was a projection of what leaving everything running to the deadline would have cost. The current month-to-date values are displayed above from the billing APIs. Azure remains lower than the projection because the gateway runs only in short bursts—built, verified and torn down again—rather than continuously.
How fresh can this be? Azure Cost Management returns same-day data, though the current day keeps trickling in for several hours. AWS Cost Explorer lags a day or more and flags its figures as Estimated. So "live cost" does not exist on either cloud. What is real time is the resource inventory, which is what actually matters for catching something left running. Any team member with the documented read-only access can run ./scripts/refresh-cost.sh; it fails closed if either cloud login or a required billing field is unavailable.
Control lesson. The first guardrails covered AWS even though credits reduce its net bill to approximately zero, while the Pay-As-You-Go Azure subscription carried the cash cost. Azure now has its own monthly budget and alerts.
Guardrails are alert-only: the cloud emails you, you act. The cheapest guardrail is still teardown · delete-stack on AWS and az group delete on Azure between work sessions.
Owner means whoever can actually act. The programme manager tracks every risk on this list and chases all of them, but only owns the ones where he holds the lever. A risk assigned to someone who cannot discharge it is not managed, it is just parked.
40.119.233.66 is kept, which is why the AWS customer gateway cgw-03e47b4fe8c6e27f1 pinned to it needs no change and the rebuild is Azure-only. About 30 minutes via REBUILD.md phase D. Owner: Bernard.This is a historical evidence record, not the hourly operational reading. Current tunnel, alarm, compute and cost summaries come from the live banner and map. Infrastructure facts below were captured after the 25 Jul rebuild; working-artefact metadata was updated on 27 Jul.
vgw-0bf466d711cc93ac4, VPN #1 vpn-065016cdb8d06a873, VPN #2 vpn-03dcb879c286f5ecd, both VPN connections, both flow logs. Verified with describe-vpn-connections, describe-customer-gateways, describe-flow-logs.SourceDestCheck=false on the strongSwan instance, GatewaySubnet 10.2.255.0/27 with nsg: None, workload subnet 10.2.1.0/24 with its NSG attached, local network gateway configured for the current AWS VPN #2 endpoint and 10.1.0.0/16.IsLogging: true with delivery at 18:48 SGT, SNS subscribed to bernard.tay@hotmail.com with a real ARN, both AWS budgets at $40 and $3, Azure budget at $25.ttl=254, against the ~69 ms recorded earlier. Within normal jitter, same path.81287885425387e67c9caba089f2944a, 280,087 bytes, 21 rendered pages; deck MD5 37b6a7f368611006dca1bcdc2ac0a462, 70,324 bytes, 13 slides with sourced speaker notes.i-074bf0cb910a99e02; replacement i-03d63eb39288c7419 served HTTP through hub.eg334s.internal at 10.1.2.225 in 244 seconds. Primary, DNS and zero-capacity steady state restored at 15:32:57 SGT. This is recoverability evidence, not full HA. Evidence and claim boundary.Assessment documents verified 28 Jul 2026. The assessment weightings, scope, report requirements and both rubrics were checked against the supplied official PDFs. The report content order shown in the brief is an example for reference, not a mandatory sequence. The current working report was rendered and checked at 21 pages, with fixed contents and footer numbering. It remains subject to team review. One infrastructure limitation remains unverified: the orphaned IAM role sits on the previous lab account, which this session has no access to, so its continued existence is reported, not observed.
Re-run this before the demo. Infrastructure claims decay. The shared ./scripts/verify-teardown.sh and ./scripts/refresh-cost.sh tools are now in this repository. Both refuse to report success if authentication or a required query fails. Use the provisional evidence register to replace historical captures after the 25 Aug rebuild.
eg334s-hub-recovery firstaz group delete --name eg334s-rg --yesaz resource list --tag Project=EG334S returns emptyAWS order matters: delete eg334s-vpn2-aws before eg334s-vpn1-awspublic, it borrows that VGW.
Note: an orphaned IAM role from the locked-down lab account (eg334s-vpn1-onprem-SsmRole-*) can't be self-deleted · flag to lab admin. IAM is global, not regional.
eg334s-vpn2-aws borrows the virtual private gateway owned by eg334s-vpn1-awspublic. Delete VPN #2 first so the owner stack can remove the gateway cleanly.
After teardown, run the authenticated verification script across both AWS regions and Azure. Retain its PASS / FAIL / ERROR transcript as the M6 completion evidence.
The historical deployment used a separate VPN #2 stack that borrowed the hub VGW by parameter, so CloudFormation did not enforce deletion order. The current published hub template can stage VPN #2 inside the same stack. Which model is live is TBD until stack-resource inventory succeeds; follow the observed model, not an assumption.
The shared verification script checks the project-named AWS resources returned by the complete inventory plus every resource in eg334s-rg, and fails closed with ERROR if authentication or any required query fails. Account-level budgets, the CloudWatch dashboard, key pairs, legacy IAM and the Cloudflare evidence site remain explicit manual/TBD controls in the ownership matrix.
Shared evidence record · technical milestones are source-controlled, not browser-local. Times in Asia/Singapore. Built for EG334S. Live status is in the banner and map above; the update stamp is below.
Le nozze di Figaro, K. 492, Act III, No. 21, Sull’aria … Che soave zeffiretto, the Letter Duet. Five interpretations. Press play above. Full tracks need an active Spotify session in this browser, otherwise Spotify serves previews. Turn this off before the assessed demo.