Received: by 10.223.185.116 with SMTP id b49csp5495664wrg; Wed, 7 Mar 2018 12:47:33 -0800 (PST) X-Google-Smtp-Source: AG47ELss1NIiYllOKKb8pQwy5Ha/a2eds8+n6ho9tUniBz9Myspr5GLJz47kz/mlwM1cwKyumxdZ X-Received: by 10.99.104.73 with SMTP id d70mr18975154pgc.107.1520455653243; Wed, 07 Mar 2018 12:47:33 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1520455653; cv=none; d=google.com; s=arc-20160816; b=mAxEnWVR6C9TiXzpdwJLTi95tee4xabBuvSPAM3wMhQBaPijfKLU2M97we85fPjUgT 5svSAY1V7xYNC+yAoZ1OErYIfCNwDY1PwYfiR2U6yr4sOLMtNvvVNhNw0vnkiGvwhOVM oCehyK2nkrB9eCkQ7moMt8SsF5UAl/JvzUHfLS9L+bO2j1g72eYoveRncMvIfxyWZDDm Iw2Eg1LHW703beNHY2+JJnEQ8TSOAJ1OUDSTko40pZdagjlKle3CjOXQi6YNhYmA30IY 7qX7kkyFTcwS/knMTC9Grd0Qw2TxQ8ApvI9E1/yr9RBZMJDG3bo+wY/3nRoLFBaBC6Bl pa4g== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:sender:mime-version:user-agent:references :in-reply-to:message-id:date:subject:cc:to:from :arc-authentication-results; bh=KQhasd6Z4uKLuhFEDan5Tfcb0qtkwC7hCuPjdOagqyU=; b=pUWXFMwJu8By5lBOWQ2K95jT02FSoV/obJm6WvISwQhVu1wQjVjBCJt6P1AFv2zlug L+j0WPNUxM0mq/kv1aidjAfONo89ksRDv1+jng51LJI22TjNzOS5HWcMPL5Eb7pe+2E4 GctBNZfFSwWqo4Vqzx3Mb9X0nLYwnHrr4kOvQA4XiM+yf/XFtc5GvM0hrFaVcY5iuulJ GVNH/qOwwOqKdhNeb0HAu7Z5cTuRPyJ8t2myk2vXO/ulOsKEbc2ZPiNVR7aiJUg/XWvF rkT2jeH1m2Br/utJ4UEAFfgsWNh3aevNaJn25+XIV6yAxTYmIwAvWp+nAW4M3ougr9dp z4Vg== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org Return-Path: Received: from vger.kernel.org (vger.kernel.org. [209.132.180.67]) by mx.google.com with ESMTP id 126si4700828pgj.673.2018.03.07.12.47.18; Wed, 07 Mar 2018 12:47:33 -0800 (PST) Received-SPF: pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) client-ip=209.132.180.67; Authentication-Results: mx.google.com; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S934596AbeCGTlH (ORCPT + 99 others); Wed, 7 Mar 2018 14:41:07 -0500 Received: from mail.linuxfoundation.org ([140.211.169.12]:40888 "EHLO mail.linuxfoundation.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S933415AbeCGTlE (ORCPT ); Wed, 7 Mar 2018 14:41:04 -0500 Received: from localhost (unknown [185.236.200.248]) by mail.linuxfoundation.org (Postfix) with ESMTPSA id 67A95FD6; Wed, 7 Mar 2018 19:41:03 +0000 (UTC) From: Greg Kroah-Hartman To: linux-kernel@vger.kernel.org Cc: Greg Kroah-Hartman , stable@vger.kernel.org, Grygorii Strashko , Ivan Khoronzhuk , "David S. Miller" Subject: [PATCH 4.15 044/122] net: ethernet: ti: cpsw: fix net watchdog timeout Date: Wed, 7 Mar 2018 11:37:36 -0800 Message-Id: <20180307191735.374113232@linuxfoundation.org> X-Mailer: git-send-email 2.16.2 In-Reply-To: <20180307191729.190879024@linuxfoundation.org> References: <20180307191729.190879024@linuxfoundation.org> User-Agent: quilt/0.65 X-stable: review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Sender: linux-kernel-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org 4.15-stable review patch. If anyone has any objections, please let me know. ------------------ From: Grygorii Strashko [ Upstream commit 62f94c2101f35cd45775df00ba09bde77580e26a ] It was discovered that simple program which indefinitely sends 200b UDP packets and runs on TI AM574x SoC (SMP) under RT Kernel triggers network watchdog timeout in TI CPSW driver (<6 hours run). The network watchdog timeout is triggered due to race between cpsw_ndo_start_xmit() and cpsw_tx_handler() [NAPI] cpsw_ndo_start_xmit() if (unlikely(!cpdma_check_free_tx_desc(txch))) { txq = netdev_get_tx_queue(ndev, q_idx); netif_tx_stop_queue(txq); ^^ as per [1] barier has to be used after set_bit() otherwise new value might not be visible to other cpus } cpsw_tx_handler() if (unlikely(netif_tx_queue_stopped(txq))) netif_tx_wake_queue(txq); and when it happens ndev TX queue became disabled forever while driver's HW TX queue is empty. Fix this, by adding smp_mb__after_atomic() after netif_tx_stop_queue() calls and double check for free TX descriptors after stopping ndev TX queue - if there are free TX descriptors wake up ndev TX queue. [1] https://www.kernel.org/doc/html/latest/core-api/atomic_ops.html Signed-off-by: Grygorii Strashko Reviewed-by: Ivan Khoronzhuk Signed-off-by: David S. Miller Signed-off-by: Greg Kroah-Hartman --- drivers/net/ethernet/ti/cpsw.c | 16 ++++++++++++++-- 1 file changed, 14 insertions(+), 2 deletions(-) --- a/drivers/net/ethernet/ti/cpsw.c +++ b/drivers/net/ethernet/ti/cpsw.c @@ -1618,6 +1618,7 @@ static netdev_tx_t cpsw_ndo_start_xmit(s q_idx = q_idx % cpsw->tx_ch_num; txch = cpsw->txv[q_idx].ch; + txq = netdev_get_tx_queue(ndev, q_idx); ret = cpsw_tx_packet_submit(priv, skb, txch); if (unlikely(ret != 0)) { cpsw_err(priv, tx_err, "desc submit failed\n"); @@ -1628,15 +1629,26 @@ static netdev_tx_t cpsw_ndo_start_xmit(s * tell the kernel to stop sending us tx frames. */ if (unlikely(!cpdma_check_free_tx_desc(txch))) { - txq = netdev_get_tx_queue(ndev, q_idx); netif_tx_stop_queue(txq); + + /* Barrier, so that stop_queue visible to other cpus */ + smp_mb__after_atomic(); + + if (cpdma_check_free_tx_desc(txch)) + netif_tx_wake_queue(txq); } return NETDEV_TX_OK; fail: ndev->stats.tx_dropped++; - txq = netdev_get_tx_queue(ndev, skb_get_queue_mapping(skb)); netif_tx_stop_queue(txq); + + /* Barrier, so that stop_queue visible to other cpus */ + smp_mb__after_atomic(); + + if (cpdma_check_free_tx_desc(txch)) + netif_tx_wake_queue(txq); + return NETDEV_TX_BUSY; }