aha/Documentation/x86 at 3c0797925f4ef9d55a32059d2af61a9c262e639d - adulau/aha - Forgejo: adulau git carryall

adulau/aha

mirror of https://github.com/adulau/aha.git synced 2024-12-27 19:26:25 +00:00

History

Andi Kleen 3c0797925f x86, mce: switch x86 machine check handler to Monarch election. On Intel platforms machine check exceptions are always broadcast to all CPUs. This patch makes the machine check handler synchronize all these machine checks, elect a Monarch to handle the event and collect the worst event from all CPUs and then process it first. This has some advantages: - When there is a truly data corrupting error the system panics as quickly as possible. This improves containment of corrupted data and makes sure the corrupted data never hits stable storage. - The panics are synchronized and do not reenter the panic code on multiple CPUs (which currently does not handle this well). - All the errors are reported. Currently it often happens that another CPU happens to do the panic first, but reports useless information (empty machine check) because the real error happened on another CPU which came in later. This is a big advantage on Nehalem where the 8 threads per CPU lead to often the wrong CPU winning the race and dumping useless information on a machine check. The problem also occurs in a less severe form on older CPUs. - The system can detect when no CPUs detected a machine check and shut down the system. This can happen when one CPU is so badly hung that that it cannot process a machine check anymore or when some external agent wants to stop the system by asserting the machine check pin. This follows Intel hardware recommendations. - This matches the recommended error model by the CPU designers. - The events can be output in true severity order - When a panic happens on another CPU it makes sure to be actually be able to process the stop IPI by enabling interrupts. The code is extremly careful to handle timeouts while waiting for other CPUs. It can't rely on the normal timing mechanisms (jiffies, ktime_get) because of its asynchronous/lockless nature, so it uses own timeouts using ndelay() and a "SPINUNIT" The timeout is configurable. By default it waits for upto one second for the other CPUs. This can be also disabled. From some informal testing AMD systems do not see to broadcast machine checks, so right now it's always disabled by default on non Intel CPUs or also on very old Intel systems. Includes fixes from Ying Huang Fixed a "ecception" in a comment (H.Seto) Moved global_nwo reset later based on suggestion from H.Seto v2: Avoid duplicate messages [ Impact: feature, fixes long standing problems. ] Signed-off-by: Andi Kleen <ak@linux.intel.com> Signed-off-by: Hidetoshi Seto <seto.hidetoshi@jp.fujitsu.com> Signed-off-by: H. Peter Anvin <hpa@zytor.com>		2009-06-03 14:45:12 -07:00
..
i386	x86: doc: move x86-generic documentation from Doc/x86/i386	2008-07-22 15:34:38 -04:00
x86_64	x86, mce: switch x86 machine check handler to Monarch election.	2009-06-03 14:45:12 -07:00
00-INDEX	documentation: move mtrr.txt to Doc/x86/ subdir	2008-07-28 14:46:49 +02:00
boot.txt	Merge branches 'x86/apic', 'x86/cpu', 'x86/fixmap', 'x86/mm', 'x86/sched', 'x86/setup-lzma', 'x86/signal' and 'x86/urgent' into x86/core	2009-03-04 02:22:31 +01:00
earlyprintk.txt	x86/doc: mini-howto for using earlyprintk=dbgp	2009-03-05 10:57:52 +01:00
mtrr.txt	documentation: move mtrr.txt to Doc/x86/ subdir	2008-07-28 14:46:49 +02:00
pat.txt	x86: PAT: pfnmap documentation update changes	2008-12-19 15:40:31 -08:00
usb-legacy-support.txt	x86: doc: move x86-generic documentation from Doc/x86/i386	2008-07-22 15:34:38 -04:00
zero-page.txt	documentation: update header file paths	2009-01-06 15:59:28 -08:00