Skip to content
All projects

CLOUD / BACKEND / DATA

live

TravelPlaner

A reliability-aware public-transport planner for Norway. It collects realtime departures from Entur across the whole Ruter network, turns them into delay distributions, and tells you when to leave and why. It runs in production on AWS behind CI/CD with health-gated deploys and automatic rollback.

  • Live on AWS
  • CI/CD with auto-rollback
  • 23 ADRs
  • 172 Python tests
Context
Personal product, built and operated end to end
Period
2026
Links
travel.mahamodul.no Private repository

The repository is private. I'm happy to walk through the code in an interview.

Stack

  • Python
  • FastAPI
  • PostgreSQL
  • Redis
  • Next.js
  • TypeScript
  • Keycloak
  • LightGBM
  • Docker
  • Terraform
  • AWS EC2
  • ECR
  • S3
  • SSM
  • GitHub Actions

Project overview

TravelPlaner answers a question normal planners skip: if I take this connection, how likely am I to be on time, and when should I leave? For each journey it shows the probability of arriving on time, the buffer you need, a leave-by time, and a plain-language explanation of every estimate.

Behind the web app is a data platform I call Norway Mobility Intelligence. It collects realtime data continuously, keeps raw snapshots immutable, aggregates delays by line, stop, weekday and hour, and serves a statistical reliability engine through a FastAPI backend. Signed-in users can save daily commutes and get a leave-by time every day.

Problem

Timetables and realtime feeds tell you where a bus is now. They don't tell you how often the 08:12 actually misses a 4-minute transfer at Majorstuen. Answering that needs history: weeks of observed departures per line, stop and hour, plus a way to combine the delays of several legs into one probability.

Requirements

  • Collect every Ruter departure through Entur's network-wide SIRI feed, plus busy hubs for other operators, without losing or rewriting raw data.
  • Search must stay fast at request time, so aggregation happens ahead of time and incrementally.
  • Every estimate must be explainable and must state its confidence when history is thin.
  • Accounts are optional, and tokens must never reach the browser.
  • Run in production on a small budget, with no manual deploy steps and a safe way back from a bad release.

Architecture

Data moves from the sources through to the clients, and the browser only ever talks to the web server.

The backend is a FastAPI modular monolith with journey, location, reliability, weather, analytics, users and commute modules. Collectors and the processor run as separate processes because they behave differently at runtime. A pure-Python reliability package holds the maths, with no web or database code, so it can be unit-tested on its own.

Technology decisions

The repository holds 23 architecture decision records. These are the ones that shaped the system most:

Modular monolith, not microservices (ADR-002)
The domain has clear boundaries, but there is one developer and modest traffic. Modules talk through in-process interfaces; only workers with a different runtime profile are separate processes.
Statistics before machine learning (ADR-005, ADR-014)
A deterministic engine with documented weights serves every estimate. A LightGBM model is only served after it beats that engine on days it has never seen and an administrator activates it, and the statistical engine stays as the automatic fallback.
Immutable raw snapshots (ADR-006)
The storage layer refuses overwrites, so any day can be reprocessed from source after a parser fix.
Backend-for-frontend sign-in (ADR-010)
The Next.js server runs the OIDC code flow with PKCE and keeps tokens in Redis. The browser only holds an opaque, httpOnly session id.
One free-plan server instead of ECS Fargate, for now (ADR-021)
The target architecture (Fargate, RDS, ElastiCache, ALB, NAT) would cost roughly $150–250 a month before any traffic. The first deployment runs the same Compose stack on one t4g.small for about $19 a month, built with Terraform so it can grow into the full design later.
Health-gated deploys (ADR-022)
A release only stays live if every service reports healthy within six minutes. Otherwise the previous version is deployed again automatically.

Implementation

The reliability engine picks the most specific delay history that has at least 20 observations. It starts with line, stop, weekday and hour, and falls back through coarser levels to documented per-mode defaults, which are always labelled low confidence. It then adjusts for realtime delay and weather, and composes the legs of a journey:

Transfer risk between two legs (packages/reliability/transfer.py)text
margin  = (aimed departure of next − aimed arrival of previous) − walking time
P(miss) = P(A_prev − B'_next > margin)

A   = adjusted arrival-delay distribution of the incoming leg
B'  = departure-delay distribution of the outgoing leg, 50% of its
      historical lateness credited (realtime delay credited in full)
  • Each leg can fail through cancellation or a missed connection. A failure costs one headway of waiting, and the final arrival is a mixture over every failure combination (computed exactly up to 6 legs).
  • The API returns probability_on_time, expected delay, p95 delay and the leave-by time, together with the explanation codes the UI turns into sentences.
  • Ruter's feed alone adds about 8,000 observed departures every few minutes, and a Ruter day is roughly 1 million rows. Detailed observations stay 14 days in PostgreSQL, then move to daily gzipped training files.
  • The ML pipeline exports training data nightly, trains weekly, compares models with the baseline on later unseen days, and supports shadow mode before activation.

Security considerations

Sign-in: tokens stay on the server
  1. 01Browser→Web (BFF)GET /auth/login
  2. 02Web (BFF)→Browser302 to Keycloak with state, nonce, PKCE S256
  3. 03Browser→KeycloakSign in
  4. 04Web (BFF)→KeycloakCode + verifier on the back channel
  5. 05Web (BFF)→RedisStore tokens; verify ID token (iss, aud, nonce)
  6. 06Web (BFF)→BrowserhttpOnly session cookie (random id only)
  7. 07Web (BFF)→APIBearer access token; API checks JWKS signature, iss, aud, exp
  • One public surface: only the website and sign-in are exposed. The API, PostgreSQL and Redis have no public host name, and the web proxy allow-lists paths and caps request bodies at 16 KB.
  • SIRI XML is parsed with defusedxml in streaming mode, with tests that confirm DTDs and entity expansion are refused.
  • Server-side ownership checks: another user's commute returns 404, the same as a missing one.
  • CSP, X-Frame-Options, rate limiting backed by Redis, and containers running as a non-root user.
  • In AWS, only ports 80 and 443 are open and there is no SSH (SSM Session Manager instead). IMDSv2 is required, and every secret is generated by Terraform and stored as an SSM SecureString.
  • CI runs pip-audit, npm audit and Trivy image scans. HIGH or CRITICAL findings fail the build.

Cloud & delivery

Every merge to main ships itself, and a failing release rolls itself back.
  • Migrations follow expand-then-contract, so a rollback never meets a schema it can't read. Every release takes a database backup first.
  • The last five images of each service stay in ECR, and deploy bundles are kept for 90 days, so any recent version can be put back with one command.
  • AWS Budgets email at 50% and 80% of the monthly limit, and a narrowly scoped role stops the server at 100% or on any real charge.

Challenges

  • Cold start: with no history, estimates fall back to documented defaults and say so (low confidence) rather than pretending to be precise.
  • Rollback safety: rolling back code does not roll back a database, which led to the expand-then-contract rule for migrations.
  • Cloud constraints: the AWS organisation policy blocked GitHub OIDC providers, so deploys use a narrowly scoped key stored only in GitHub's encrypted production environment, with a documented path back to OIDC.
  • Memory: the full stack uses about 1.6 GB, so the server runs with memory caps per container and 2 GB of swap.

Results

travel.mahamodul.no
Live
on AWS since Sep 2026
per month
~$19
estimated, versus $150–250 for the target design
health gate
≤ 6 min
before a release stays live
Python test functions
172
unit, contract, API and integration

Figures from the repository's deployment docs and test suite.

The machine-learning stage is built up to activation. A fair evaluation against the statistical engine needs about three weeks of collected history, so the model is not yet serving estimates.

What I learned

  • Writing decisions down (ADRs) made later trade-offs faster, because the original constraints were still visible.
  • A cheap deployment can still be a careful one: health gates, backups before migrations and a budget kill switch matter more than the size of the server.
  • Shipping a statistical baseline first gave the ML work an honest benchmark and a safe fallback.

Source