Received: by 2002:a25:ad19:0:0:0:0:0 with SMTP id y25csp8935397ybi; Wed, 10 Jul 2019 01:56:53 -0700 (PDT) X-Google-Smtp-Source: APXvYqyMX+in/Ql1fH8WXJYAsFrLmcmFfh/OGtBsyHa9dwZ7CVkF+EyH2zSKG6WK5IeF13FBkUZH X-Received: by 2002:a17:90a:1aa4:: with SMTP id p33mr5705302pjp.27.1562749013456; Wed, 10 Jul 2019 01:56:53 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1562749013; cv=none; d=google.com; s=arc-20160816; b=AsbCaaF8X5CtIPC1anBcpLP7UTTuiZBFdDV/Ft06r7Y4sQcw9+bO7ttPc/fPokhbxJ qf4wvQet7iaewWGcwJHfKRQq6kxwn66Jez+sLRQ1JpEAeRWBPSJ/ET8gb49f2KP6A1BQ +p+5Qr9pfl34jjPp6pmtk8XIfpSLOu19LUzpWaSP9sQgaECin2Xk6cFCUrPiXyjVNa0n PjXUnUZDdEIxTQ39qrIacGvwW0GLtu8qu6UAkSeOkwFv2cKspgVFyUc7Y8w4/+M7QPae j4hreEyIDOdV+LrKAY+24Lac3j3VsHJKlJRHkLES2Xss5y0m09Hl7dA5FyzMt6iQNbU7 7Qkw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:sender:content-transfer-encoding:cc:to:subject :message-id:date:from:in-reply-to:references:mime-version :dkim-signature; bh=6btS87boNqG1W/9Ztpa2IpNskVJOh6JYgS9t3inEIQI=; b=nYxkqmPqANM0Em/ArC56AdpnZvzw5ZDo5HYnth+0cJb8DEoqvmGNIKI8CV4+aFrozb 7CSsldhS5uzpZO+i9OIbqiVSy6AdKzcIGRq1ckWTtWAeoF+M1VlMCCoJDG4Z0TGBs1Pf htvwuQRWvF9NeG78KnDbBUNUS1/9DvYuz2hjPADC4c4TEVrZ6oqKmH8VG2YgOkVg83Nx MaOE/4xRdrnR1NF0+4lcn/cIdbgj8wqOM/yCM9ab7IQrf0pKE8NSMRW3D4mxvwLvDc+k y0TVOu6vdVQrPcEaRGg0xeEUA3Fa3O/Fe2tTE4tP409+hncGSs3TH1q0owLXk/4Rxtux JlQA== ARC-Authentication-Results: i=1; mx.google.com; dkim=pass header.i=@endlessm-com.20150623.gappssmtp.com header.s=20150623 header.b=KMlYx9EO; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org Return-Path: Received: from vger.kernel.org (vger.kernel.org. [209.132.180.67]) by mx.google.com with ESMTP id u9si1617955pgb.148.2019.07.10.01.56.36; Wed, 10 Jul 2019 01:56:53 -0700 (PDT) Received-SPF: pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) client-ip=209.132.180.67; Authentication-Results: mx.google.com; dkim=pass header.i=@endlessm-com.20150623.gappssmtp.com header.s=20150623 header.b=KMlYx9EO; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1727591AbfGJIy6 (ORCPT + 99 others); Wed, 10 Jul 2019 04:54:58 -0400 Received: from mail-io1-f67.google.com ([209.85.166.67]:46094 "EHLO mail-io1-f67.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1727588AbfGJIy5 (ORCPT ); Wed, 10 Jul 2019 04:54:57 -0400 Received: by mail-io1-f67.google.com with SMTP id i10so2975770iol.13 for ; Wed, 10 Jul 2019 01:54:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=endlessm-com.20150623.gappssmtp.com; s=20150623; h=mime-version:references:in-reply-to:from:date:message-id:subject:to :cc:content-transfer-encoding; bh=6btS87boNqG1W/9Ztpa2IpNskVJOh6JYgS9t3inEIQI=; b=KMlYx9EOkoY8Qy/PDR4ey635CnLrzEeFzu2f9fIL5n1fifS9gS8a1BkIkyU31MnjiM Y3nz80EheAksIOa13APe2tRja5415pH7L5g5lEo41gMjbs9pP3JgbBNq+h9gTE8zSjBv rmYvAck4xvjmpWCOMDGXtQqInmhZtJ0dulXOs/yqW9QDzmu578ULxBUex3izQ/Vudti+ leN8v/A+HVewkXLG8ycgF6WvQ/9J/Z+2I9myzqHeA5cXITUrWiiuay3Tjsvh+0n8kdOp j7LEOhW7cuhaZBsZGIclq/X0Huql21CWkgL4DdRYBtAMEkgaP1Pq2PxBFR1Pwvtqs5rH AVsw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:mime-version:references:in-reply-to:from:date :message-id:subject:to:cc:content-transfer-encoding; bh=6btS87boNqG1W/9Ztpa2IpNskVJOh6JYgS9t3inEIQI=; b=cViN1wVr5qoXNZ0w+w5geLMfJ8rbMLWtI+4s1XG3di75CDoDdUvgYSwWDGZI9+dmqr HQ6W5LPUo36GwxxUyF48CsAFtcaSArr1RyKzuyb1OIJIMHLHVWicoJ38h1mINOky/7Bm VxD6bkojwKOgZeY+5Zg6lfMaadnks66vKEcMeparwP2wnWKyn2VVuwh/5FHobSWjoV1/ oXZfuO97xPZALs0YfiWCefpx6cH3Vd0wckRORV4oW/lRROmkkrjQRVKUaPQ9ssVzeHt4 kHXa860sn7VDDweW76VKIm2+FNlollUOLBV0Q2o6+NjostR2MzFUw9xprE/EajwiQYQd d+7A== X-Gm-Message-State: APjAAAWAhFe5tVJvZOGe7UghbBCD+fXey9cFJp1+AngOpU+6YIjh/l/H /JX139NMfSaF49VhVOUAiHtRFY1xcILN13kiaUuWtw== X-Received: by 2002:a02:b10b:: with SMTP id r11mr6018114jah.140.1562748896251; Wed, 10 Jul 2019 01:54:56 -0700 (PDT) MIME-Version: 1.0 References: <20190708063252.4756-1-jian-hong@endlessm.com> <20190709102059.7036-1-jian-hong@endlessm.com> In-Reply-To: From: Jian-Hong Pan Date: Wed, 10 Jul 2019 16:54:19 +0800 Message-ID: Subject: Re: [PATCH v2 1/2] rtw88: pci: Rearrange the memory usage for skb in RX ISR To: Tony Chuang Cc: Kalle Valo , "David S . Miller" , Larry Finger , David Laight , "linux-wireless@vger.kernel.org" , "netdev@vger.kernel.org" , "linux-kernel@vger.kernel.org" , "linux@endlessm.com" , Daniel Drake , "stable@vger.kernel.org" Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Sender: linux-kernel-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Tony Chuang =E6=96=BC 2019=E5=B9=B47=E6=9C=8810=E6= =97=A5 =E9=80=B1=E4=B8=89 =E4=B8=8B=E5=8D=884:37=E5=AF=AB=E9=81=93=EF=BC=9A > > > Subject: [PATCH v2 1/2] rtw88: pci: Rearrange the memory usage for skb = in > > RX ISR > > > > Testing with RTL8822BE hardware, when available memory is low, we > > frequently see a kernel panic and system freeze. > > > > First, rtw_pci_rx_isr encounters a memory allocation failure (trimmed): > > > > rx routine starvation > > WARNING: CPU: 7 PID: 9871 at drivers/net/wireless/realtek/rtw88/pci.c:8= 22 > > rtw_pci_rx_isr.constprop.25+0x35a/0x370 [rtwpci] > > [ 2356.580313] RIP: 0010:rtw_pci_rx_isr.constprop.25+0x35a/0x370 [rtwpc= i] > > > > Then we see a variety of different error conditions and kernel panics, > > such as this one (trimmed): > > > > rtw_pci 0000:02:00.0: pci bus timeout, check dma status > > skbuff: skb_over_panic: text:00000000091b6e66 len:415 put:415 > > head:00000000d2880c6f data:000000007a02b1ea tail:0x1df end:0xc0 > > dev: > > ------------[ cut here ]------------ > > kernel BUG at net/core/skbuff.c:105! > > invalid opcode: 0000 [#1] SMP NOPTI > > RIP: 0010:skb_panic+0x43/0x45 > > > > When skb allocation fails and the "rx routine starvation" is hit, the > > function returns immediately without updating the RX ring. At this > > point, the RX ring may continue referencing an old skb which was alread= y > > handed off to ieee80211_rx_irqsafe(). When it comes to be used again, > > bad things happen. > > > > This patch allocates a new, data-sized skb first in RX ISR. After > > copying the data in, we pass it to the upper layers. However, if skb > > allocation fails, we effectively drop the frame. In both cases, the > > original, full size ring skb is reused. > > > > In addition, to fixing the kernel crash, the RX routine should now > > generally behave better under low memory conditions. > > > > Buglink: https://bugzilla.kernel.org/show_bug.cgi?id=3D204053 > > Signed-off-by: Jian-Hong Pan > > Cc: > > --- > > drivers/net/wireless/realtek/rtw88/pci.c | 49 +++++++++++------------- > > 1 file changed, 22 insertions(+), 27 deletions(-) > > > > diff --git a/drivers/net/wireless/realtek/rtw88/pci.c > > b/drivers/net/wireless/realtek/rtw88/pci.c > > index cfe05ba7280d..e9fe3ad896c8 100644 > > --- a/drivers/net/wireless/realtek/rtw88/pci.c > > +++ b/drivers/net/wireless/realtek/rtw88/pci.c > > @@ -763,6 +763,7 @@ static void rtw_pci_rx_isr(struct rtw_dev *rtwdev, > > struct rtw_pci *rtwpci, > > u32 pkt_offset; > > u32 pkt_desc_sz =3D chip->rx_pkt_desc_sz; > > u32 buf_desc_sz =3D chip->rx_buf_desc_sz; > > + u32 new_len; > > u8 *rx_desc; > > dma_addr_t dma; > > > > @@ -790,40 +791,34 @@ static void rtw_pci_rx_isr(struct rtw_dev *rtwdev= , > > struct rtw_pci *rtwpci, > > pkt_offset =3D pkt_desc_sz + pkt_stat.drv_info_sz + > > pkt_stat.shift; > > > > - if (pkt_stat.is_c2h) { > > - /* keep rx_desc, halmac needs it */ > > - skb_put(skb, pkt_stat.pkt_len + pkt_offset); > > + /* discard current skb if the new skb cannot be allocated= as a > > + * new one in rx ring later > > + */ > > + new_len =3D pkt_stat.pkt_len + pkt_offset; > > + new =3D dev_alloc_skb(new_len); > > + if (WARN_ONCE(!new, "rx routine starvation\n")) > > + goto next_rp; > > + > > + /* put the DMA data including rx_desc from phy to new skb= */ > > + skb_put_data(new, skb->data, new_len); > > > > - /* pass offset for further operation */ > > - *((u32 *)skb->cb) =3D pkt_offset; > > - skb_queue_tail(&rtwdev->c2h_queue, skb); > > + if (pkt_stat.is_c2h) { > > + /* pass rx_desc & offset for further operation *= / > > + *((u32 *)new->cb) =3D pkt_offset; > > + skb_queue_tail(&rtwdev->c2h_queue, new); > > ieee80211_queue_work(rtwdev->hw, &rtwdev->c2h_wor= k); > > } else { > > - /* remove rx_desc, maybe use skb_pull? */ > > - skb_put(skb, pkt_stat.pkt_len); > > - skb_reserve(skb, pkt_offset); > > - > > - /* alloc a smaller skb to mac80211 */ > > - new =3D dev_alloc_skb(pkt_stat.pkt_len); > > - if (!new) { > > - new =3D skb; > > - } else { > > - skb_put_data(new, skb->data, skb->len); > > - dev_kfree_skb_any(skb); > > - } > > - /* TODO: merge into rx.c */ > > - rtw_rx_stats(rtwdev, pkt_stat.vif, skb); > > + /* remove rx_desc */ > > + skb_pull(new, pkt_offset); > > + > > + rtw_rx_stats(rtwdev, pkt_stat.vif, new); > > memcpy(new->cb, &rx_status, sizeof(rx_status)); > > ieee80211_rx_irqsafe(rtwdev->hw, new); > > } > > > > - /* skb delivered to mac80211, alloc a new one in rx ring = */ > > - new =3D dev_alloc_skb(RTK_PCI_RX_BUF_SIZE); > > - if (WARN(!new, "rx routine starvation\n")) > > - return; > > - > > - ring->buf[cur_rp] =3D new; > > - rtw_pci_reset_rx_desc(rtwdev, new, ring, cur_rp, buf_desc= _sz); > > +next_rp: > > + /* new skb delivered to mac80211, re-enable original skb = DMA */ > > + rtw_pci_reset_rx_desc(rtwdev, skb, ring, cur_rp, buf_desc= _sz); > > > > /* host read next element in ring */ > > if (++cur_rp >=3D ring->r.len) > > -- > > 2.22.0 > > Now it looks good to me. Thanks. > > Acked-by: Yan-Hsuan Chuang > > Yan-Hsuan Uh! Thanks for your ack. But I just sent version 3 patches (including [PATCH v3 2/2] rtw88: pci: Use DMA sync instead of remapping in RX ISR) by following Christoph's comment. [1] Could you please also review the 2 patches of version 3? Thank you. [1]: https://lkml.org/lkml/2019/7/9/507 Jian-Hong Pan