Received: by 2002:a05:6a10:22f:0:0:0:0 with SMTP id 15csp2425438pxk; Sat, 3 Oct 2020 22:02:39 -0700 (PDT) X-Google-Smtp-Source: ABdhPJzKakrXD8N6brhiAclSpErleaSFmrSAHzLyBTc0UYYlDlyai/Kh66w/i/XbA8kdRD9CSVYr X-Received: by 2002:a17:906:6855:: with SMTP id a21mr9177934ejs.289.1601787758980; Sat, 03 Oct 2020 22:02:38 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1601787758; cv=none; d=google.com; s=arc-20160816; b=YUZyiSEhQCzp7Ir6VgGqyWK0c4+Xr5qoyQxpt+/hkcHb5d7jyk6+AgQKq66oW5ISQW /YrUk7YLFho9oXs6ujvukv2Ipn4+MwbrvuMYp0fPvVAHuaiER2fTSG+dFsTa7HJ8IcZ6 dcXc3NLrN4knuIaccbBdcfp5JbjPhenC9Xmslr9NUZEPYOFVVZAVUCNbdLX4+fG42+v5 Rk7HUrv46c5gqYxP5U+/RmBaQEDFomQneHnR988FM1905RmoBYhYNJYipGNBs0i7THbu B980wbeWfFyXuOhwCOhpWywTxOaFn9ZSDz/FjVdE6lb/tFe8VfUsi0AnP45+awd5V0gO qVHQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:user-agent:in-reply-to:content-disposition :mime-version:references:message-id:subject:cc:to:from:date :ironport-sdr:ironport-sdr; bh=pMeiq9tI1RlHNziE7YPunFz+rlrnfDbn79vEKX3hVhs=; b=0SsFf5BsS4Ll6QGxCVZHTRmjD7MKBRYHcTvG9MMAVndcvc+OQXSGHgQrM7EJMr4k0l uNomyxTrDv914dSALih4cHBJM9g7KVpp6H4dnJE4DS0/cQRWG9S/jcke6/yACrkdYT2s 7OEOcUKuopfzWO4cQhwT1ytCoKVAVHqNUD/1PS/OIHwMZLueC/jTbRA+Df5MlNR9uDZn h71eDv1EW0K/2kHhTKbu26S4TccW6tqb6hYeWHqClpIl7I9X7UNKAXs6oTt8O/MjHpsT SS1tmveBIZiBGsozDrFfD2XEwb4/U1OIdCZBiGhacFpxXtmkQG0yvCGTNXhHSpOCysss HvPw== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=intel.com Return-Path: Received: from vger.kernel.org (vger.kernel.org. [23.128.96.18]) by mx.google.com with ESMTP id n15si4805169edy.300.2020.10.03.22.01.47; Sat, 03 Oct 2020 22:02:38 -0700 (PDT) Received-SPF: pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) client-ip=23.128.96.18; Authentication-Results: mx.google.com; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=intel.com Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1725825AbgJDE5v (ORCPT + 99 others); Sun, 4 Oct 2020 00:57:51 -0400 Received: from mga01.intel.com ([192.55.52.88]:60183 "EHLO mga01.intel.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1725818AbgJDE5v (ORCPT ); Sun, 4 Oct 2020 00:57:51 -0400 IronPort-SDR: xsJTYiHMU9isfRf2SZdE1jwUUhtgnF3UIROIT03V2xI1dTi4IJKmnw/pBHKKsQxhVtCZMDjOkh DIJHf5af6pLQ== X-IronPort-AV: E=McAfee;i="6000,8403,9763"; a="181387079" X-IronPort-AV: E=Sophos;i="5.77,334,1596524400"; d="scan'208";a="181387079" X-Amp-Result: SKIPPED(no attachment in message) X-Amp-File-Uploaded: False Received: from orsmga005.jf.intel.com ([10.7.209.41]) by fmsmga101.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 03 Oct 2020 21:57:50 -0700 IronPort-SDR: 1m64lsAECejvgOfT786z1x8HjoeT8NuY7DMdZ6dxJj68S/dd4q6aoFI3R85h+1em24QBB7Fadh rBV9lTPdiMIA== X-IronPort-AV: E=Sophos;i="5.77,334,1596524400"; d="scan'208";a="517996501" Received: from araj-mobl1.jf.intel.com ([10.251.22.42]) by orsmga005-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 03 Oct 2020 21:57:48 -0700 Date: Sat, 3 Oct 2020 21:57:47 -0700 From: "Raj, Ashok" To: Ethan Zhao Cc: bhelgaas@google.com, oohall@gmail.com, ruscur@russell.cc, lukas@wunner.de, andriy.shevchenko@linux.intel.com, stuart.w.hayes@gmail.com, mr.nuke.me@gmail.com, mika.westerberg@linux.intel.com, linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, sathyanarayanan.kuppuswamy@intel.com, xerces.zhao@gmail.com, Ashok Raj Subject: Re: [PATCH v7 0/5] Fix DPC hotplug race and enhance error handling Message-ID: <20201004045745.GA3207@araj-mobl1.jf.intel.com> References: <20201003075514.32935-1-haifeng.zhao@intel.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20201003075514.32935-1-haifeng.zhao@intel.com> User-Agent: Mutt/1.9.1 (2017-09-22) Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi Ethan On Sat, Oct 03, 2020 at 03:55:09AM -0400, Ethan Zhao wrote: > Hi,folks, > > This simple patch set fixed some serious security issues found when DPC > error injection and NVMe SSD hotplug brute force test were doing -- race > condition between DPC handler and pciehp, AER interrupt handlers, caused > system hang and system with DPC feature couldn't recover to normal > working state as expected (NVMe instance lost, mount operation hang, > race PCIe access caused uncorrectable errors reported alternatively etc). I think maybe picking from other commit messages to make this description in cover letter bit clear. The fundamental premise is that when due to error conditions when events are processed by both DPC handler and hotplug handling of DLLSC both operating on the same device object ends up with crashes. > > With this patch set applied, stable 5.9-rc6 on ICS (Ice Lake SP platform, > see > https://en.wikichip.org/wiki/intel/microarchitectures/ice_lake_(server)) > > could pass the PCIe Gen4 NVMe SSD brute force hotplug test with any time > interval between hot-remove and plug-in operation tens of times without > any errors occur and system works normal. > > With this patch set applied, system with DPC feature could recover from > NON-FATAL and FATAL errors injection test and works as expected. > > System works smoothly when errors happen while hotplug is doing, no > uncorrectable errors found. > > Brute DPC error injection script: > > for i in {0..100} > do > setpci -s 64:02.0 0x196.w=000a > setpci -s 65:00.0 0x04.w=0544 > mount /dev/nvme0n1p1 /root/nvme > sleep 1 > done > > Other details see every commits description part. > > This patch set could be applied to stable 5.9-rc6/rc7 directly. > > Help to review and test. > > v2: changed according to review by Andy Shevchenko. > v3: changed patch 4/5 to simpler coding. > v4: move function pci_wait_port_outdpc() to DPC driver and its > declaration to pci.h. (tip from Christoph Hellwig ). > v5: fix building issue reported by lkp@intel.com with some config. > v6: move patch[3/5] as the first patch according to Lukas's suggestion. > and rewrite the comment part of patch[3/5]. > v7: change the patch[4/5], based on Bjorn's code and truth table. > change the patch[5/5] about the debug output information. > > Thanks, > Ethan > > > Ethan Zhao (5): > PCI/ERR: get device before call device driver to avoid NULL pointer > dereference > PCI/DPC: define a function to check and wait till port finish DPC > handling > PCI: pciehp: check and wait port status out of DPC before handling > DLLSC and PDC > PCI: only return true when dev io state is really changed > PCI/ERR: don't mix io state not changed and no driver together > > drivers/pci/hotplug/pciehp_hpc.c | 4 ++- > drivers/pci/pci.h | 55 +++++++++++++------------------- > drivers/pci/pcie/dpc.c | 27 ++++++++++++++++ > drivers/pci/pcie/err.c | 18 +++++++++-- > 4 files changed, 68 insertions(+), 36 deletions(-) > > > base-commit: a1b8638ba1320e6684aa98233c15255eb803fac7 > -- > 2.18.4 >