| Meetings | Tuesdays & Thursdays, 12:00–1:30 pm, AGH 214 |
|---|---|
| Instructor | Josh Fried — [email protected] |
| Office hours | Tuesdays 2–3 pm, Levine 604 |
| Piazza | Piazza — announcements, logistics, and paper questions before each session |
| Canvas | Canvas — lecture slides, project guidelines, submissions (warm-up, proposal, report), and grades |
| Testbed | CloudLab |
Overview
This graduate-level seminar explores recent research in operating systems, focusing on the software infrastructure of the datacenter. The course provides a technical tour of how modern datacenter stacks are architected and examines the challenges inherent in managing resources at warehouse scale. We read, present, and discuss research papers spanning multicore OS architecture, kernel-bypass I/O, OS abstractions, microsecond-scale scheduling, hardware offload, memory management, virtualization and isolation, cluster management, datacenter storage, and machine-learning systems. Students complete a semester-long systems project.
Course Structure
Most sessions cover two related papers. Read the assigned paper(s) and post to Piazza, by 9:00 am on the day of class, a short take on each plus a few questions — things you found unclear, surprising, or worth arguing about. These are ungraded; they exist so the session starts from your questions. They count toward participation and can earn up to 2 bonus points. Most sessions begin with a short in-class quiz on the assigned reading (the lowest seven quiz scores are dropped). Each student signs up to lead roughly two sessions (a ~15–20 minute overview of each paper, then guided discussion); sign-ups open Thu Aug 27 and the first student-led session is Thu Sep 10. On most paper sessions two other students take short rotating roles — Skeptic (three minutes making the case against the paper, including at least one objection about the evaluation) and Historian (three minutes on what this replaced and what happened next). There will be an early warm-up assignment on CloudLab to get you familiar with the computing resources available to you there (assigned Thu Sep 3, due Thu Sep 17).
Prerequisites
A graduate operating systems course (or equivalent), comfort with systems programming in C/C++ and Linux, and networking basics (e.g., an undergraduate networking course). Familiarity with computer architecture is helpful but not required.
Research project
Open-ended and semester-long. Work individually or in pairs; groups are capped at two, and a pair is expected to deliver a correspondingly more substantial project. Strong projects identify a real inefficiency or limitation in a proposed research system, prototype a change, and evaluate it experimentally; reproducing and extending a recent result is also welcome. Using your own graduate research is fine, provided it is related to the material in this course; please discuss scope with the instructor early. Project guidelines are posted Tue Sep 1 — a separate document describing the kinds of project that work in this course, what each has to deliver, and a list of seed ideas you are free to take, adapt, or ignore. The seeds are a starting point for brainstorming, not well-defined projects. Read it before you write your proposal. Class time on Tue Sep 1 covers CloudLab and previews several research systems Josh has built — Shenango/Caladan, Junction, Spice — that might make good substrates for a semester's work. Milestones: proposal (~1–2 pages) due Fri Sep 25; mid-semester checkpoint the week of Oct 19–23; final presentations Tue Dec 1 & Thu Dec 3; a required 20-minute final check-in per group on the reading days, Tue Dec 8 or Wed Dec 9, to walk through your results before submitting; final report (~6 pages, workshop-paper style) due Tue Dec 15. The report is submitted together with your code and a script that regenerates every figure in it from your raw data.
Grading
| Research project | 40% |
| Quizzes (in class, on the readings) | 20% |
| Class participation | 20% |
| Discussion leadership & presentation | 15% |
| Warm-up assignment (CloudLab) | 5% |
| Piazza paper questions | ungraded — up to +2 bonus |
Schedule
The schedule below is subject to change — topics and readings may shift as the semester progresses; check back for updates.
Paper questions and comments are posted to Piazza by 9:00 am on the day of class; most sessions begin with a short quiz on the reading.
Papers tagged industry describe deployed production
systems; academic papers are research proposals.
| Date | Topic | Reading | Presenter | Notes |
|---|---|---|---|---|
| Part 0 — Introduction | ||||
| Tue Aug 25 | Introduction, course overview, and how to read a paper | optional: The Datacenter as a Computer (ch. 1–2, skim) (Barroso, Hölzle & Ranganathan, 4th ed. 2025) industry optional: Twenty Five Years of Warehouse-Scale Computing (in memory of Luiz Barroso) (Ranganathan & Hölzle, IEEE Micro 2024) industry | Josh Fried | |
| Thu Aug 27 | The hardware–OS gap | Josh Fried | Leader sign-ups open | |
| Tue Sep 1 | Sources of tail latency + Platforms for research projects (second half) | Tales of the Tail: Hardware, OS, and Application-level Sources of Tail Latency (Li et al., SoCC 2014) academic | Josh Fried | Project guidelines posted |
| Thu Sep 3 | Multi-tenancy and oversubscription | Resource Central: Understanding and Predicting Workloads for Improved Resource Management in Large Cloud Platforms (Cortez et al., SOSP 2017) industry optional: PerfIso: Performance Isolation for Commercial Latency-Sensitive Services (Iorgulescu et al., ATC 2018) industry | Josh Fried | Warm-up assigned |
| Part 1 — The OS and Modern Hardware | ||||
| Tue Sep 8 | Multicore OS scalability | The Multikernel: A New OS Architecture for Scalable Multicore Systems (Baumann et al., SOSP 2009) academic optional: An Analysis of Performance Evolution of Linux's Core Operations (Ren et al., SOSP 2019) academic | Josh Fried | |
| Thu Sep 10 | Kernel-bypass I/O | IX: A Protected Dataplane Operating System for High Throughput and Low Latency (Belay et al., OSDI 2014) academic | — | |
| Tue Sep 15 | The kernel storage stack and fast SSDs | Rearchitecting Linux Storage Stack for µs Latency and High Throughput (blk-switch) (Hwang et al., OSDI 2021) academic | — | |
| Thu Sep 17 | Host networking in production | Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization (Dalton et al., NSDI 2018) industry | — | Warm-up due |
| Tue Sep 22 | RPCs and RDMA | A Cloud-Optimized Transport Protocol for Elastic and Scalable HPC (SRD) (Shalev et al., IEEE Micro 2020) industry | — | |
| Part 2 — Interfaces and Abstractions | ||||
| Thu Sep 24 | Datapath abstractions | optional: POSIX Abstractions in Modern Operating Systems: The Old, the New, and the Missing (Atlidakis et al., EuroSys 2016) academic | — | Project proposal due Fri Sep 25 |
| Part 3 — Scheduling & Resource Sharing | ||||
| Tue Sep 29 | Core scheduling | Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads (Ousterhout et al., NSDI 2019) academic ghOSt: Fast & Flexible User-Space Delegation of Linux Scheduling (Humphries et al., SOSP 2021) industry | — | |
| Thu Oct 1 | No class — Fall Term Break | |||
| Tue Oct 6 | Preemption and time slicing | Skyloft: A General High-Efficient Scheduling Framework in User Space (Jia et al., SOSP 2024) academic | — | Mid-semester feedback survey (in class) |
| Thu Oct 8 | Datacenter transport | Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities (Montazeri et al., SIGCOMM 2018) academic | — | |
| Part 4 — The Datacenter Tax & Hardware Offload | ||||
| Tue Oct 13 | The datacenter tax | Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at Hyperscale (Sriraman et al., ASPLOS 2020) industry optional: Retrospective: Profiling a Warehouse-Scale Computer (2 pp.) (Kanev et al., ISCA@50, 2023) industry optional: Profiling Hyperscale Big Data Processing (Spanner, BigTable, BigQuery) (Gonzalez et al., ISCA 2023) industry | — | |
| Thu Oct 15 | Network offload | — | ||
| Tue Oct 20 | Offloading infrastructure to DPUs | — | Checkpoint this week | |
| Part 5 — Isolation | ||||
| Thu Oct 22 | Performance isolation | — | ||
| Tue Oct 27 | MicroVMs and containers | Firecracker: Lightweight Virtualization for Serverless Applications (Agache et al., NSDI 2020) industry Blending Containers and Virtual Machines: A Study of Firecracker and gVisor (Anjali et al., VEE 2020) academic | — | |
| Thu Oct 29 | Guest lecture: serverless cold starts | Rethinking Process Snapshots for Near-Warm Serverless Cold Starts (Spice) (Holmes et al., OSDI 2026) academic | Ben Holmes (MIT CSAIL) | |
| Part 6 — Memory | ||||
| Tue Nov 3 | Project hacking day (Election Day) — in-class project work and office hours | — | — | |
| Thu Nov 5 | Memory offloading and pooling | optional: AIFM: High-Performance, Application-Integrated Far Memory (Ruan et al., OSDI 2020) academic | — | |
| Part 7 — The Cluster | ||||
| Tue Nov 10 | Cluster management | optional: Twine: A Unified Cluster Management System for Shared Infrastructure (Tang et al., OSDI 2020) industry | — | |
| Thu Nov 12 | Datacenter storage | optional: What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block Store (Zhang et al., FAST 2024) industry | — | |
| Part 8 — Machine Learning Systems | ||||
| Tue Nov 17 | LLM serving | Orca: A Distributed Serving System for Transformer-Based Generative Models (Yu et al., OSDI 2022) academic Efficient Memory Management for LLM Serving with PagedAttention (vLLM) (Kwon et al., SOSP 2023) academic optional: Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving (Qin et al., FAST 2025) industry | — | |
| Thu Nov 19 | GPU networking | — | ||
| Part 9 — Project Presentations | ||||
| Tue Nov 24 | Project hacking day — in-class project work and office hours (meets on Penn's Thursday schedule) | — | — | |
| Thu Nov 26 | No class — Thanksgiving Break | |||
| Tue Dec 1 | Project presentations | — | project teams | |
| Thu Dec 3 | Project presentations | — | project teams | Final check-in Dec 8–9 · Final report + code due Tue Dec 15 |
Policies
- Attendance. This is a discussion course; showing up prepared is the assignment. Tell the instructor in advance if you must miss. More than two unexcused absences may affect the participation grade.
- Participation. You participate by contributing to the class's thinking: speaking in discussion, taking a rotating role when it's your turn, posting a question or a disagreement on Piazza, or asking a good question of a classmate who is presenting. The instructor is looking for evidence that you read the paper and formed a view about it, not for airtime — three careful contributions across a session beat fifteen reflexive ones, and a student who is quiet in the room but consistently sharp on Piazza will do fine. If you find speaking in a group difficult, say so early and we will find a version of this that works for you.
- Piazza posts. Ungraded and voluntary, but the session is built out of them — if the thread is empty, the discussion is worse for everyone. Post by 9:00 am so the discussion leader can use what you wrote. Quizzes cannot be made up except for excused absences.
- Mid-semester feedback. On Tue Oct 6 the instructor will take five minutes of class for an anonymous survey on how the course is going: reading load, quizzes, discussion, project support. Results and any resulting changes are reported back at the following session.
- AI tools. AI tools can be great to use when reading papers — to answer questions, unpack notation, and bring in background a paper assumes. Their use for this is encouraged. Students remain ultimately responsible for learning the material: quizzes and in-class discussion are unassisted, and the Piazza posts exist to find out what you think. AI tools may also be used to assist in projects, depending on the goal of the project; please discuss this with the instructor before relying on them, and document in your report what you used and how.
- Academic integrity. All submitted work must be your own (or your group's), per Penn's Code of Academic Integrity. The instructor may ask you to walk through any work you submit and explain how it was produced.
- Accessibility and accommodations. Penn provides reasonable accommodations to students with disabilities and temporary medical conditions. Accommodations are arranged through Disability Services at the Weingarten Center (215-573-9235), which reviews requests individually and confidentially; you do not need to disclose anything about your condition to the instructor. Start there — the review can take up to four weeks, so begin early if you have not registered before. Once your accommodation letter is issued, send it to the instructor and it will be implemented. If something about the format of this course is a barrier and you would rather just say so directly, that is welcome too.
- Health and wellbeing. Graduate coursework is not worth your health. If you are struggling — with this course or with anything else — Wellness at Penn offers free, confidential counseling to all Penn students; you can reach a counselor at any hour at 215-746-WELL (9355). Medical care is at 215-746-3535. In an emergency on campus, call PennComm at 215-573-3333. Come talk to the instructor if a deadline in this course is the problem; a date is much easier to move than a crisis is to undo.
- Religious and secular observance. Penn's Policy on Secular and Religious Holidays applies. If an observance conflicts with a class meeting, a quiz, or a deadline, tell the instructor within the first two weeks of the semester — even if the exact date is not fixed yet — and an alternative will be arranged. Election Day, Tue Nov 3, is a project work session with no reading, no quiz, and nothing due.
Further Reading
Additional papers for context and to browse for project ideas.
Workloads & measurement
- The Scalable Commutativity Rule (SOSP 2013)
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload (ATC 2020)
- Latency Lags Bandwidth (CACM 2004)
- AI and Memory Wall (IEEE Micro 2024)
- Server Architecture From Enterprise to Post-Moore (IEEE Micro 2024)
Storage datapath
- Asynchronous I/O Stack: A Low-latency Kernel I/O Stack for Ultra-Low Latency SSDs (ATC 2019)
- ReFlex: Remote Flash ≈ Local Flash (ASPLOS 2017)
Abstractions
- Unikernels: Library Operating Systems for the Cloud (ASPLOS 2013)
Performance isolation & colocation
- Iron: Isolating Network-based CPU in Container Environments (NSDI 2018)
- Heracles: Improving Resource Efficiency at Scale (ISCA 2015)
- PicNIC: Predictable Virtualized NIC (SIGCOMM 2019)
- CPI²: CPU Performance Isolation for Shared Compute Clusters (EuroSys 2013)
- PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services (ASPLOS 2019)
Scheduling & dataplanes
- ZygOS: Achieving Low Tail Latency for Microsecond-scale Networked Tasks (SOSP 2017)
- Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal Scheduling (Concord) (SOSP 2023)
- Arachne: Core-Aware Thread Management (OSDI 2018)
Transport & load balancing
- Swift: Delay is Simple and Effective for Congestion Control (SIGCOMM 2020)
- R2P2: Making RPCs First-Class Datacenter Citizens (ATC 2019)
- Load Is Not What You Should Balance: Introducing Prequal (NSDI 2024)
RDMA & RPCs
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters (SIGCOMM 2020)
- FaRM: Fast Remote Memory (NSDI 2014)
- Design Guidelines for High Performance RDMA Systems (ATC 2016)
- PRISM: Rethinking the RDMA Interface for Distributed Systems (SOSP 2021)
NICs, DPUs & accelerators
- The nanoPU: A Nanosecond Network Stack for Datacenters (OSDI 2021)
- Offloading Distributed Applications onto SmartNICs using iPipe (SIGCOMM 2019)
- From Luna to Solar: The Evolutions of the Compute-to-Storage Networks in Alibaba Cloud (SIGCOMM 2022)
- Achelous: Enabling Programmability, Elasticity, and Reliability in Hyperscale Cloud Networks (SIGCOMM 2023)
- Triton: A Flexible Hardware Offloading Architecture for Accelerating Apsara vSwitch in Alibaba Cloud (SIGCOMM 2024)
- In-Datacenter Performance Analysis of a Tensor Processing Unit (ISCA 2017)
- A Cloud-Scale Acceleration Architecture (Catapult) (MICRO 2016)
Memory
- TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory (ASPLOS 2023)
- Nu: Achieving Microsecond-Scale Resource Fungibility with Logical Processes (NSDI 2023)
- LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation (OSDI 2018)
- Clio: A Hardware-Software Co-Designed Disaggregated Memory System (ASPLOS 2022)
- Software-Defined Far Memory in Warehouse-Scale Computers (ASPLOS 2019)
- Don't Shoot Down TLB Shootdowns! (EuroSys 2020)
- Mitosis: Transparently Self-Replicating Page-Tables (ASPLOS 2020)
Virtualization & isolation
- Xen and the Art of Virtualization (SOSP 2003)
- Hey, You, Get Off of My Cloud: Exploring Information Leakage in Third-Party Compute Clouds (CCS 2009)
- Confidential VMs Explained: An Empirical Analysis of AMD SEV-SNP and Intel TDX (SIGMETRICS 2025)
- Security and Performance in the Delegated User-level Virtualization (DuVisor) (OSDI 2023)
- Keystone: An Open Framework for Architecting TEEs (EuroSys 2020)
- Occlum: Secure and Efficient Multitasking Inside an SGX Enclave (ASPLOS 2020)
- Faasm: Lightweight Isolation for Efficient Stateful Serverless Computing (ATC 2020)
- Unikraft: Fast, Specialized Unikernels the Easy Way (EuroSys 2021)
- SOCK: Rapid Task Provisioning with Serverless-Optimized Containers (ATC 2018)
Cluster management & operations
- Take It to the Limit: Peak Prediction-driven Resource Overcommitment in Datacenters (EuroSys 2021)
- Twine: A Unified Cluster Management System for Shared Infrastructure (OSDI 2020)
- Protean: VM Allocation Service at Scale (OSDI 2020)
- Autopilot: Workload Autoscaling at Google (EuroSys 2020)
- Borg: the Next Generation (EuroSys 2020)
- Sparrow: Distributed, Low Latency Scheduling (SOSP 2013)
- Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google TR 2010)
- Canopy: An End-to-End Performance Tracing and Analysis System (SOSP 2017)
Storage systems
- Flat Datacenter Storage (OSDI 2012)
- What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block Store (FAST 2024)
LLM systems
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (OSDI 2024)
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve (OSDI 2024)
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving (FAST 2025)
- Splitwise: Efficient Generative LLM Inference Using Phase Splitting (ISCA 2024)
- Alibaba HPN: A Data Center Network for LLM Training (SIGCOMM 2024)
- TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches (NSDI 2023)
Sustainability & power
- Data Center Power and Energy Management: Past, Present, and Future (IEEE Micro 2024)
- Beyond Efficiency: Scaling AI Sustainably (IEEE Micro 2024)
- Carbon-Aware Computing for Datacenters (IEEE TPS 2023)
- Treehouse: A Case for Carbon-Aware Datacenter Software (HotCarbon 2022)