Operating Systems — Lecture 02
A C program that prints one line:
$ sudo perf stat -e instructions ./hello
Hello World!
1,015,131 instructions
A freestanding assembly program that writes 14 bytes and exits:
$ sudo perf stat -e instructions ./hello_nostd
Hello, World!
108,573 instructions
The program wrote a dozen instructions. The machine ran a hundred thousand.
Common features. Every program starts, gets memory, reads files, talks to the network. Written once, below every application, reached through one interface.
Isolation. Many programs, one machine, none entitled to trust the others. Something has to keep them apart — and can, because the hardware gives it a privilege they lack.
The operating system is a provider of services — which is why it has an interface.
It is an enforcer of isolation — which is why that interface is a boundary, not a function call.
Like a library — but not linked into your program, and shared by every program at once.
A normal library trusts its caller. The kernel must not: its callers are every program on the machine.
A unikernel is an OS compiled down to one application — the smallest an OS can be.
$ ls -lh c-hello_qemu-x86_64
-rwxr-xr-x 241K c-hello_qemu-x86_64
$ nm c-hello_qemu-x86_64.dbg | wc -l
855
240 KB and 800 symbols, to print one line. “Run a program on a machine” is a big job.
Load rax with the call number, arguments in rdi, rsi, rdx, r10, r8, r9, execute syscall. Result in rax.
$ strace ./read_write_syscall
write(1, "Gimme message: ", 15) = 15
read(0, "hello there\n", 64) = 12
write(1, "hello there\n", 12) = 12
exit_group(0) = ?
The trace is the program. Nothing sits between it and the kernel.
| Family | Examples | Lecture |
|---|---|---|
| processes and threads | fork(), execve(), clone() |
06–08 |
| memory | brk(), mmap(), munmap() |
03–05 |
| file I/O | open(), read(), write() |
09–10 |
| network I/O | socket(), connect(), send() |
11 |
| IPC | pipes, signals, shared memory | 10, 12 |
A few hundred in Linux: man 2 syscalls.
Programs call printf(), fopen(), malloc(). libc turns those into system calls — when it has to.
getpid() — one system callstrlen() never; malloc() only when it runs outprintf() — a write(), or nothing, depending on the bufferlibc decides when to cross. That is what part 04 exploits.
User mode: cannot touch device registers, page tables, interrupts. Kernel mode: can.
$ ./cli
first
Segmentation fault (core dumped)
cli disables interrupts; mov rax, cr3 reads the page-table base. Both are ring-0 only.
There is no check for cli in the kernel. The CPU faults on it at ring 3.
A software check can have a bug. This cannot.
kernel/user = a hardware mode: which instructions run now. root/non-root = a software identity: what the kernel will agree to do.
Root still runs in user mode. Root still faults on cli. The axes are orthogonal.
Ten million one-byte writes to /dev/null (kernel work: nil), against the same loop as a function call:
$ ./make_syscalls
time passed 703224 microseconds
$ ./make_libcalls
time passed 9360 microseconds
~70x. /dev/null does nothing — almost all of it is the crossing itself.
That is not waste. It is the price of the check from part 03.
$ ./fwrite_buffered
time passed 65846 microseconds
$ ./fwrite_unbuffered
time passed 1508197 microseconds
Buffer on: one write per few thousand bytes. Buffer off: one per byte.
A few kilobytes of memory, a twenty-fold speedup, identical output. This is why libc buffers.
sendfile() moves a file to a socket without the bytes entering user space:
$ ./server_sendfile ~1800 us
$ ./server_write ~5900 us
server_write: read + send per chunk, two copies. sendfile: one call, no bounce through user space.
One idea, two ends: batch the crossings, or remove them.
No correct answer. Linux is monolithic; seL4 is a microkernel and flies aircraft.
A hypervisor does for operating systems what an OS does for applications.
More privileged than the kernel. Multiplexes the hardware. Isolates its guests. Each guest kernel believes it owns the machine.
OS : processes :: hypervisor : operating systems
The software stack from lecture 01, extended one layer down.
The operating system has a dual role: provider of services, enforcer of isolation.
Isolation needs a boundary. The boundary is entered through system calls, it costs, and OS types and virtualization are where you put it.
my_syscall(), libc-style helpers on top, strace to checkFull write-up, demos and references: the session README.md.