Resume

Software Development Engineer at Amazon, Berlin, with ~4 years building and operating distributed backend systems on AWS that serve millions of requests worldwide. I specialize in event-driven architecture, high availability, and reliability engineering — root-causing the incidents others work around and re-architecting fragile systems so they stop paging. I also work hands-on with AI coding agents like Claude and Kiro — grounding them with service context and reusable skills to automate on-call triage and engineering workflows, and contributing to OpenHands, an open-source AI agent platform where I've built LLM-integrated features. M.Sc. in High Integrity Systems from Frankfurt UAS, with prior research in real-time systems at DLR.

Experience

Amazon - Software Development Engineer

- Present

Designing, delivering, and operating distributed backend services on Amazon's retail platform at scale — from zero-to-one system builds to cross-org incident root-causing and reliability engineering.
  • Architected a unified notification platform from scratch on AWS Step Functions, consolidating two legacy services after evaluating three candidate architectures — cut new notification-type onboarding from ~1 month to ~1 week (a partner team shipped a new type end-to-end in 7 days).
  • Root-caused 4 months of recurring Sev-2 incidents (23+ total) in distributed DynamoDB-stream processing — where propagation lag spiked from minutes to 29+ hours across 4 regions — and led the migration to an event-driven serverless (Lambda) architecture; zero incidents since.
  • Diagnosed an org-wide cache race condition generating 9,000+ errors/hour on a service handling millions of requests — an issue multiple teams had only worked around; root-caused and deployed a worldwide fix in a single day, raising service availability from 97.45% to 99.999% and sustaining it through the rest of the year.
  • Cut a cross-region compliance-migration outage window from ~6 hours to under 15 minutes (24x) by designing a pre-staged deployment across 3 teams in 3 time zones — zero lost events and zero manual intervention; adopted as the template for later migrations.
  • Traced a 50% drop in a leadership-reported conversion metric through the full analytics pipeline to 36 bot accounts producing 87% of traffic with zero submissions — proved no customer regression and delivered both a self-service query and a permanent pipeline fix.
  • Apply AI coding agents fluently across day-to-day engineering — writing and reviewing code, drafting docs, and accelerating on-call investigation — and author service-specific agent skills and context so the agent stays grounded and accurate.
  • Built an AI onboarding skill for one of my team's services so away teams could self-onboard and resolve questions about how the service works without support; and orchestrated a multi-team campaign where scheduled AI jobs tracked responses across tickets — updating docs, recording deadlines, marking completion, and flagging only the tickets needing my attention.

German Aerospace Center (DLR) - Master Thesis Researcher

-

Research on response time analysis of real-time task chains for satellite software at DLR's Institute for Software Technology.
  • Developed a novel worst-case response-time analysis for sporadic DAG tasks on multi-core processors under preemptive global fixed-priority scheduling — proving that assigning intra-task priorities to DAG subtasks tightens the safe upper bound on response time.
  • Designed an algorithm that controls subtask execution order to derive that bound, and modeled sporadic tasks without collapsing them to periodic — reducing analytical pessimism in inter-task interference.
  • Validated in Python across hundreds of randomly generated DAGs (up to 150 nodes / 8 threads, via Erdös–Rényi and Nested Fork-Join generators), showing the method outperforms state-of-the-art approaches; published at DLR elib.

Amazon - SDE Intern

-

End-to-end feature delivery on Amazon Retail Mobile — design, implementation, and worldwide rollout.
  • Designed and shipped the redesigned customer-review image gallery (Java, Spring, Angular) — a masonry-layout, lazy-loaded page replacing the legacy experience — and owned its worldwide rollout; drove a 78% increase in image views, validated via A/B testing.

Software AG - Working Student — R&D

-

Built internal tooling and optimized production services at a global enterprise software company (€800M+ revenue), alongside M.Sc. studies.
  • Optimized Java RESTful services — cutting production errors by 30% — and built an internal university-relations platform with Spring Boot, Angular, and Keycloak SSO.

NTT Data FA Insurance Systems - Software Engineer

-

Delivered enterprise insurance applications serving 5,000+ agents across multiple clients in the APAC region.
  • Built a workload-distribution feature that replaced manual job assignment with efficient round-robin distribution across client employees — initially for one insurance client, then promoted into the core product for all clients after its impact; used by 5,000+ agents.
  • Automated WebLogic deployment via Java/WLST scripting — reduced deployment times by 90%.

Education

Skills

AI & Agentic Engineering
AI Coding Agents (Claude, Kiro)Agent Skills & Context Engineering
Backend
JavaSpring / Spring BootMicroservicesRESTful APIsEvent-Driven ArchitectureNode.js
Cloud & Infrastructure
AWS (CDK, Lambda, DynamoDB, SQS, CloudWatch)DockerCI/CD PipelinesInfrastructure as CodeHigh Availability & ScalabilityObservability & MonitoringStream Processing (Kinesis, DynamoDB Streams)DynamoDBKubernetes
Data & Research
PythonReal-Time Systems
Databases
SQLDynamoDBMySQLPostgreSQL
Languages
JavaPythonTypeScriptJavaScriptSQLC
System Design
MicroservicesEvent-Driven ArchitectureDistributed SystemsHigh Availability & ScalabilityObservability & MonitoringReal-Time SystemsStream Processing (Kinesis, DynamoDB Streams)Incident Response & On-Call
Web Development
RESTful APIsTypeScriptJavaScriptNode.jsReactAngularNext.js

References available upon request. Get in touch →