LinuxLists.cc - [PATCH] [0/21] HWPOISON: Intro

2009-09-11 18:52:04

Subject: [PATCH] [0/21] HWPOISON: Intro

This the version of hwpoison I intend to submit for 2.6.32.

Only some very minor fixes compared to the last version posted.
I integrated one patch from Fengguang that has been reviewed
separately.

Passes the hwpoison specific parts of the mce-test test suite
(git://git.kernel.org/pub/scm/utils/cpu/mce/mce-test.git)

Also available as git tree from
git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-mce-2.6.git hwpoison

-Andi

Signed-off-by: Andi Kleen <[email protected]>

---

Upcoming Intel CPUs have support for recovering from some memory errors
(``MCA recovery''). This requires the OS to declare a page "poisoned",
kill the processes associated with it and avoid using it in the future.

This patchkit implements the necessary infrastructure in the VM.

To quote the overview comment:

* High level machine check handler. Handles pages reported by the
* hardware as being corrupted usually due to a 2bit ECC memory or cache
* failure.
*
* This focusses on pages detected as corrupted in the background.
* When the current CPU tries to consume corruption the currently
* running process can just be killed directly instead. This implies
* that if the error cannot be handled for some reason it's safe to
* just ignore it because no corruption has been consumed yet. Instead
* when that happens another machine check will happen.
*
* Handles page cache pages in various states. The tricky part
* here is that we can access any page asynchronous to other VM
* users, because memory failures could happen anytime and anywhere,
* possibly violating some of their assumptions. This is why this code
* has to be extremely careful. Generally it tries to use normal locking
* rules, as in get the standard locks, even if that means the
* error handling takes potentially a long time.
*
* Some of the operations here are somewhat inefficient and have non
* linear algorithmic complexity, because the data structures have not
* been optimized for this case. This is in particular the case
* for the mapping from a vma to a process. Since this case is expected
* to be rare we hope we can get away with this.

The code consists of a the high level handler in mm/memory-failure.c,
a new page poison bit and various checks in the VM to handle poisoned
pages.

The main target right now is KVM guests, but it works for all kinds
of applications.

For the KVM use there was need for a new signal type so that
KVM can inject the machine check into the guest with the proper
address. This in theory allows other applications to handle
memory failures too. The expection is that near all applications
won't do that, but some very specialized ones might.

This is not fully complete yet, in particular there are still ways
to access poison through various ways (crash dump, /proc/kcore etc.)
that need to be plugged too.

-Andi

2009-09-11 18:48:31

Subject: [PATCH] [0/21] HWPOISON: Intro

Subject: [PATCH] [1/21] HWPOISON: Add page flag for poisoned pages

Subject: [PATCH] [2/21] HWPOISON: Export some rmap vma locking to outside world

Subject: [PATCH] [3/21] HWPOISON: Add support for poison swap entries v2

Subject: [PATCH] [4/21] HWPOISON: Add new SIGBUS error codes for hardware poison signals

Subject: [PATCH] [5/21] HWPOISON: Add basic support for poisoned pages in fault handler v3

Subject: [PATCH] [6/21] HWPOISON: Add various poison checks in mm/memory.c v2

Subject: [PATCH] [7/21] HWPOISON: x86: Add VM_FAULT_HWPOISON handling to x86 page fault handler v2

Subject: [PATCH] [8/21] HWPOISON: Use bitmask/action code for try_to_unmap behaviour

Subject: [PATCH] [9/21] HWPOISON: Handle hardware poisoned pages in try_to_unmap

Subject: [PATCH] [10/21] HWPOISON: check and isolate corrupted free pages v2

Subject: [PATCH] [11/21] HWPOISON: Refactor truncate to allow direct truncating of page v2

Subject: [PATCH] [12/21] HWPOISON: Add invalidate_inode_page

Subject: [PATCH] [13/21] HWPOISON: Define a new error_remove_page address space op for async truncation

Subject: [PATCH] [14/21] HWPOISON: shmem: call set_page_dirty() with locked page

Subject: [PATCH] [15/21] HWPOISON: Add PR_MCE_KILL prctl to control early kill behaviour per process

Subject: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: [PATCH] [17/21] HWPOISON: Enable .remove_error_page for migration aware file systems

Subject: [PATCH] [18/21] HWPOISON: Enable error_remove_page for NFS

Subject: [PATCH] [19/21] HWPOISON: Add madvise() based injector for hardware poisoned pages v4

Subject: [PATCH] [20/21] HWPOISON: Add simple debugfs interface to inject hwpoison on arbitary PFNs

Subject: [PATCH] [21/21] HWPOISON: Enable error_remove_page on btrfs

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: Re: [PATCH] [14/21] HWPOISON: shmem: call set_page_dirty() with locked page

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: Re: [PATCH] [14/21] HWPOISON: shmem: call set_page_dirty() with locked page

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7

Subject: Re: [PATCH] [16/21] HWPOISON: The high level memory error handler in the VM v7