What is thrashing and how do you deal with it?
Thrashing is when a system spends more time paging than doing useful work. It happens when the total working set of active processes exceeds physical memory, so pages are evicted and immediately faulted back in. CPU utilisation collapses and disk I/O saturates.
Signs include high page-fault rates, low CPU utilisation, long run queues and heavy swap activity.
Causes and fixes:
- Too many processes for the available RAM: reduce the degree of multiprogramming or add memory.
- A poor replacement policy: use a good approximation of LRU such as clock, and keep frequently reused pages resident.
- A hot loop touching more data than fits in cache or memory: improve locality of reference.
- Aggressive swapping: tune swap behaviour, though on SSDs swapping can still destroy latency.
The working-set model keeps each process's recently used pages resident and refuses to admit a process that would push the total past available frames, which is a standard admission-control remedy.