Received: by 2002:ab2:f03:0:b0:1ef:ffd0:ce49 with SMTP id i3csp96277lqf; Tue, 26 Mar 2024 15:56:28 -0700 (PDT) X-Forwarded-Encrypted: i=3; AJvYcCXrZ1ZGAnmLWXQHpAUjzjGBEQxQg535texOrS4TvR5t+lW4jzxBXDVHFnmRyGiQpsqeupaSdKAK97/Xtalsh8jqyXwBGQuV7a11A1+goA== X-Google-Smtp-Source: AGHT+IE0LMcm0kfTOki8xOspU7Px5se5nWA01yrj7L2EQ/OPu64bddd88owAddKZfyx2JK2vz1Es X-Received: by 2002:a05:6808:14cb:b0:3c3:c2d6:e12f with SMTP id f11-20020a05680814cb00b003c3c2d6e12fmr12962628oiw.8.1711493787910; Tue, 26 Mar 2024 15:56:27 -0700 (PDT) ARC-Seal: i=2; a=rsa-sha256; t=1711493787; cv=pass; d=google.com; s=arc-20160816; b=Xf00QXhfJCYRKAyw35zDI1+A1IOxTVNAuV7O6UvIPZyEIPBFoisieqnYgi10m7HG0V tAOEL0eTQ2QUM7RAzdoLG/KruhYFOtzmYQnuXwzIuLOWyM7tglsZ2F9LdN2oI+be9UxW HKEZrAnj+uuzwNl6soMxEzLTWj4INKiWKQ2bxiZnWYkSFcJ7jr4nHgszRylmJf1NBKZC YNSJf5jWThr2UvAkjvQeBSX11fQKl4iOtOtp/aB7nTyRGYVIfYnj/bTyY9btK7kOEQTw j3/9Djc0l1W7bIqTmmhfthv7IAYZDX0o1f8PBvGc90OGy2Y8NUxseayz8FFizYN8qVmK NvoA== ARC-Message-Signature: i=2; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=cc:to:from:subject:message-id:references:mime-version :list-unsubscribe:list-subscribe:list-id:precedence:in-reply-to:date :dkim-signature; bh=E4NPaM5natLoBukffWYJGpyczDTSfX6IYb3b0cFCB0U=; fh=KlwnPP8hZpaQvI7mwOJB/riZ88pe4m44lISINs26S/0=; b=CDRS4J+WG6NnUX/meDZjX1/ThjWAhcgbrPX4BJepIaj30cI7O/UtFXqh+hs7+d3xHY yuQWViXGCCKFh59D/a4LpuFsN7ZPlqub+0uAEBaIdDhT7oFWaBxdtS4ioR6OO4qQggnL Vs8CvE9JLKVK51FRlIuLwC4Bu2kzyW/+TE565ZKZX37RD8xIvtiGm5DT8LZYb5EHn7os DyaS1Kbh6FC0F6CmMn23hwJkwn7bBUSmuqYHCpiQZKu6o9bXMB5GrBmRPLdbMk3WCSHY yl2y0yT+HWOIrbygZ2s/jIee9fmQbVw2RUbfNb2sTVFGBg5FWHZ64ywObUvCZb09K3S1 ypmQ==; dara=google.com ARC-Authentication-Results: i=2; mx.google.com; dkim=pass header.i=@google.com header.s=20230601 header.b=BLgQuy5q; arc=pass (i=1 spf=pass spfdomain=flex--almasrymina.bounces.google.com dkim=pass dkdomain=google.com dmarc=pass fromdomain=google.com); spf=pass (google.com: domain of linux-kernel+bounces-120114-linux.lists.archive=gmail.com@vger.kernel.org designates 147.75.199.223 as permitted sender) smtp.mailfrom="linux-kernel+bounces-120114-linux.lists.archive=gmail.com@vger.kernel.org"; dmarc=pass (p=REJECT sp=REJECT dis=NONE) header.from=google.com Return-Path: Received: from ny.mirrors.kernel.org (ny.mirrors.kernel.org. [147.75.199.223]) by mx.google.com with ESMTPS id v20-20020ac85794000000b004313afc772bsi8405259qta.160.2024.03.26.15.56.27 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 26 Mar 2024 15:56:27 -0700 (PDT) Received-SPF: pass (google.com: domain of linux-kernel+bounces-120114-linux.lists.archive=gmail.com@vger.kernel.org designates 147.75.199.223 as permitted sender) client-ip=147.75.199.223; Authentication-Results: mx.google.com; dkim=pass header.i=@google.com header.s=20230601 header.b=BLgQuy5q; arc=pass (i=1 spf=pass spfdomain=flex--almasrymina.bounces.google.com dkim=pass dkdomain=google.com dmarc=pass fromdomain=google.com); spf=pass (google.com: domain of linux-kernel+bounces-120114-linux.lists.archive=gmail.com@vger.kernel.org designates 147.75.199.223 as permitted sender) smtp.mailfrom="linux-kernel+bounces-120114-linux.lists.archive=gmail.com@vger.kernel.org"; dmarc=pass (p=REJECT sp=REJECT dis=NONE) header.from=google.com Received: from smtp.subspace.kernel.org (wormhole.subspace.kernel.org [52.25.139.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ny.mirrors.kernel.org (Postfix) with ESMTPS id 8EE981C66E24 for ; Tue, 26 Mar 2024 22:56:27 +0000 (UTC) Received: from localhost.localdomain (localhost.localdomain [127.0.0.1]) by smtp.subspace.kernel.org (Postfix) with ESMTP id BF11E144304; Tue, 26 Mar 2024 22:51:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="BLgQuy5q" Received: from mail-yb1-f201.google.com (mail-yb1-f201.google.com [209.85.219.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3438C13F45B for ; Tue, 26 Mar 2024 22:51:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1711493489; cv=none; b=GrQ1+E1bzBcbDALxHkxx/D8BdVRJ3ng3V/snjcgT2Ap1MjF6HZO857mqHeL7xwRXCVlZ0ri7yk+S82Q8AmkJ9XakRlljl8lW/fgJ8SsWaR0DzymH1iSEp8MfRAtqe66hpgrziK1ta6Ps0L01Hf0ITc9B7PxJJIX4F+bPixKt5Ww= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1711493489; c=relaxed/simple; bh=N4Rbcma555zxMX+UFE3uew9bgJ8/V4XQ9VXsz+6nrUE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=nBygxB6vWdlxHkaN6GkC5q+TmoCoJTsmNepY2cTK51+c6tetXFT0dNxM+GpVlajxSnB0OnAaMYb5w3doDVP1vFiDadAnSPxGr3wcKe8iNT8Ch03FvtEV6cqn5gKa28dNTONpmwQhoA1gllhQw6Z1Fkxy7c9wr61Zh9U4TxjxT3U= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--almasrymina.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=BLgQuy5q; arc=none smtp.client-ip=209.85.219.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--almasrymina.bounces.google.com Received: by mail-yb1-f201.google.com with SMTP id 3f1490d57ef6-dcd94cc48a1so9688414276.3 for ; Tue, 26 Mar 2024 15:51:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1711493481; x=1712098281; darn=vger.kernel.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=E4NPaM5natLoBukffWYJGpyczDTSfX6IYb3b0cFCB0U=; b=BLgQuy5qC0ZNVdF4byLKn5Ul1uPCLbSlMzs3uKBtmTrQ1djfTgUvjdzI4k2c5HGD0c acIo0/j5pXo0MsToSF4G/uf32bJUD8no6n2YGn92JSkyWE0Y1bmqk7dEnr2GnIhFr2wv Z4cR5Kiwe7bxsgXIOOX/yhGhWMP/Rp0DdUymeP1mkgAU0ic1F961qPZEmM5bsOJB8+KT /Un9VoXTd1IaplkGTWU1S/TraJp8Rh2P4YjkZMWoUb+GkNsx0z6SN/j1/LpG+CuZnFK2 IyotE6RL0OYeB4ZxVmugTIQWng9RNnIH1ndasVZHy0pQRvzSIz2S09AuttzHy9xdK1N6 35PQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1711493481; x=1712098281; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=E4NPaM5natLoBukffWYJGpyczDTSfX6IYb3b0cFCB0U=; b=r2IZ/i9RGKkUW6N4hGdKFwCozkrJOSUFLqbKW6fIw18ChZL370MMCkpxhbDTNiup0J l3HIEtpAHcJY5z07U1HhmbRYVgD6ru3lAjuv5iHXv+xkwXAj413MMmEh9wgB/DPPSjf1 EDD5ngeJkMRH/H6C+LKMoXiQAc4ovNOFpBnd7J5QMz1NREqum55GmlKlPHkQ0mrUHo7Y 38UocFB9Ma/9Rz2v6VbcrWFK2/d+sfEBb0C+MP5Kl/xtsSBxbNzL8Rdr7QOvEpoBAL2L P6JHwpo07M6FtcA46A7wJ1ihSQ2RH/6OLH3U9/HdXlm3g4E4U+pGZe0w0P4+47S0C2Yx Ktxg== X-Forwarded-Encrypted: i=1; AJvYcCX24lsXHVtzUWuUxyMbXX2ObSrmj+lDyiRfR37v6nY/e+SVBD/qEj1dsq4exlTsvZKUxGwUIC+fU6DWszAuBlYI++rcQcIHjnIX/lY7 X-Gm-Message-State: AOJu0YxN8S/8m0B+GyAoUFK92M9s/W8x+VAdtbM51KGFvYTjsLQ2KApG FUKmzoU9xFWchq4uRrHt2ZbNgm1oCBjWYvBC6aafp0Ncb5biiwmxc12hu0TPNG97e6kBS1SS+76 Islm3X+M60tCTEt6RpgLZQw== X-Received: from almasrymina.svl.corp.google.com ([2620:15c:2c4:200:c51e:bdd0:7cc8:695c]) (user=almasrymina job=sendgmr) by 2002:a05:6902:260f:b0:dc7:68b5:4f3d with SMTP id dw15-20020a056902260f00b00dc768b54f3dmr3456842ybb.11.1711493481430; Tue, 26 Mar 2024 15:51:21 -0700 (PDT) Date: Tue, 26 Mar 2024 15:50:43 -0700 In-Reply-To: <20240326225048.785801-1-almasrymina@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20240326225048.785801-1-almasrymina@google.com> X-Mailer: git-send-email 2.44.0.396.g6e790dbe36-goog Message-ID: <20240326225048.785801-13-almasrymina@google.com> Subject: [RFC PATCH net-next v7 12/14] net: add SO_DEVMEM_DONTNEED setsockopt to release RX frags From: Mina Almasry To: netdev@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-alpha@vger.kernel.org, linux-mips@vger.kernel.org, linux-parisc@vger.kernel.org, sparclinux@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arch@vger.kernel.org, bpf@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org Cc: Mina Almasry , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Jonathan Corbet , Richard Henderson , Ivan Kokshaysky , Matt Turner , Thomas Bogendoerfer , "James E.J. Bottomley" , Helge Deller , Andreas Larsson , Jesper Dangaard Brouer , Ilias Apalodimas , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Arnd Bergmann , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Martin KaFai Lau , Eduard Zingerman , Song Liu , Yonghong Song , John Fastabend , KP Singh , Stanislav Fomichev , Hao Luo , Jiri Olsa , Steffen Klassert , Herbert Xu , David Ahern , Willem de Bruijn , Shuah Khan , Sumit Semwal , "=?UTF-8?q?Christian=20K=C3=B6nig?=" , Pavel Begunkov , David Wei , Jason Gunthorpe , Yunsheng Lin , Shailend Chand , Harshitha Ramamurthy , Shakeel Butt , Jeroen de Borst , Praveen Kaligineedi , Willem de Bruijn , Kaiyuan Zhang Content-Type: text/plain; charset="UTF-8" Add an interface for the user to notify the kernel that it is done reading the devmem dmabuf frags returned as cmsg. The kernel will drop the reference on the frags to make them available for reuse. Signed-off-by: Willem de Bruijn Signed-off-by: Kaiyuan Zhang Signed-off-by: Mina Almasry --- v7: - Updated SO_DEVMEM_* uapi to use the next available entry (Arnd). v6: - Squash in locking optimizations from edumazet@google.com. With his changes we lock the xarray once per sock_devmem_dontneed operation rather than once per frag. Changes in v1: - devmemtoken -> dmabuf_token (David). - Use napi_pp_put_page() for refcounting (Yunsheng). - Fix build error with missing socket options on other asms. --- arch/alpha/include/uapi/asm/socket.h | 1 + arch/mips/include/uapi/asm/socket.h | 1 + arch/parisc/include/uapi/asm/socket.h | 1 + arch/sparc/include/uapi/asm/socket.h | 1 + include/uapi/asm-generic/socket.h | 1 + include/uapi/linux/uio.h | 4 ++ net/core/sock.c | 61 +++++++++++++++++++++++++++ 7 files changed, 70 insertions(+) diff --git a/arch/alpha/include/uapi/asm/socket.h b/arch/alpha/include/uapi/asm/socket.h index ef4656a41058..251b73c5481e 100644 --- a/arch/alpha/include/uapi/asm/socket.h +++ b/arch/alpha/include/uapi/asm/socket.h @@ -144,6 +144,7 @@ #define SCM_DEVMEM_LINEAR SO_DEVMEM_LINEAR #define SO_DEVMEM_DMABUF 79 #define SCM_DEVMEM_DMABUF SO_DEVMEM_DMABUF +#define SO_DEVMEM_DONTNEED 80 #if !defined(__KERNEL__) diff --git a/arch/mips/include/uapi/asm/socket.h b/arch/mips/include/uapi/asm/socket.h index 414807d55e33..8ab7582291ab 100644 --- a/arch/mips/include/uapi/asm/socket.h +++ b/arch/mips/include/uapi/asm/socket.h @@ -155,6 +155,7 @@ #define SCM_DEVMEM_LINEAR SO_DEVMEM_LINEAR #define SO_DEVMEM_DMABUF 79 #define SCM_DEVMEM_DMABUF SO_DEVMEM_DMABUF +#define SO_DEVMEM_DONTNEED 80 #if !defined(__KERNEL__) diff --git a/arch/parisc/include/uapi/asm/socket.h b/arch/parisc/include/uapi/asm/socket.h index 2b817efd4544..38fc0b188e08 100644 --- a/arch/parisc/include/uapi/asm/socket.h +++ b/arch/parisc/include/uapi/asm/socket.h @@ -136,6 +136,7 @@ #define SCM_DEVMEM_LINEAR SO_DEVMEM_LINEAR #define SO_DEVMEM_DMABUF 79 #define SCM_DEVMEM_DMABUF SO_DEVMEM_DMABUF +#define SO_DEVMEM_DONTNEED 80 #if !defined(__KERNEL__) diff --git a/arch/sparc/include/uapi/asm/socket.h b/arch/sparc/include/uapi/asm/socket.h index 00248fc68977..57084ed2f3c4 100644 --- a/arch/sparc/include/uapi/asm/socket.h +++ b/arch/sparc/include/uapi/asm/socket.h @@ -137,6 +137,7 @@ #define SCM_DEVMEM_LINEAR SO_DEVMEM_LINEAR #define SO_DEVMEM_DMABUF 0x0058 #define SCM_DEVMEM_DMABUF SO_DEVMEM_DMABUF +#define SO_DEVMEM_DONTNEED 0x0059 #if !defined(__KERNEL__) diff --git a/include/uapi/asm-generic/socket.h b/include/uapi/asm-generic/socket.h index 25a2f5255f52..1acb77780f10 100644 --- a/include/uapi/asm-generic/socket.h +++ b/include/uapi/asm-generic/socket.h @@ -135,6 +135,7 @@ #define SO_PASSPIDFD 76 #define SO_PEERPIDFD 77 +#define SO_DEVMEM_DONTNEED 97 #define SO_DEVMEM_LINEAR 98 #define SCM_DEVMEM_LINEAR SO_DEVMEM_LINEAR #define SO_DEVMEM_DMABUF 99 diff --git a/include/uapi/linux/uio.h b/include/uapi/linux/uio.h index 3a22ddae376a..d17f8fcd93ec 100644 --- a/include/uapi/linux/uio.h +++ b/include/uapi/linux/uio.h @@ -33,6 +33,10 @@ struct dmabuf_cmsg { */ }; +struct dmabuf_token { + __u32 token_start; + __u32 token_count; +}; /* * UIO_MAXIOV shall be at least 16 1003.1g (5.4.1.1) */ diff --git a/net/core/sock.c b/net/core/sock.c index 43bf3818c19e..b589610cbe4a 100644 --- a/net/core/sock.c +++ b/net/core/sock.c @@ -1049,6 +1049,63 @@ static int sock_reserve_memory(struct sock *sk, int bytes) return 0; } +#ifdef CONFIG_PAGE_POOL +static noinline_for_stack int +sock_devmem_dontneed(struct sock *sk, sockptr_t optval, unsigned int optlen) +{ + unsigned int num_tokens, i, j, k, netmem_num = 0; + struct dmabuf_token *tokens; + netmem_ref netmems[16]; + int ret; + + if (sk->sk_type != SOCK_STREAM || sk->sk_protocol != IPPROTO_TCP) + return -EBADF; + + if (optlen % sizeof(struct dmabuf_token) || + optlen > sizeof(*tokens) * 128) + return -EINVAL; + + tokens = kvmalloc_array(128, sizeof(*tokens), GFP_KERNEL); + if (!tokens) + return -ENOMEM; + + num_tokens = optlen / sizeof(struct dmabuf_token); + if (copy_from_sockptr(tokens, optval, optlen)) + return -EFAULT; + + ret = 0; + + xa_lock_bh(&sk->sk_user_frags); + for (i = 0; i < num_tokens; i++) { + for (j = 0; j < tokens[i].token_count; j++) { + netmem_ref netmem = (__force netmem_ref)__xa_erase( + &sk->sk_user_frags, tokens[i].token_start + j); + + if (netmem && + !WARN_ON_ONCE(!netmem_is_net_iov(netmem))) { + netmems[netmem_num++] = netmem; + if (netmem_num == ARRAY_SIZE(netmems)) { + xa_unlock_bh(&sk->sk_user_frags); + for (k = 0; k < netmem_num; k++) + WARN_ON_ONCE(!napi_pp_put_page(netmems[k], + false)); + netmem_num = 0; + xa_lock_bh(&sk->sk_user_frags); + } + ret++; + } + } + } + + xa_unlock_bh(&sk->sk_user_frags); + for (k = 0; k < netmem_num; k++) + WARN_ON_ONCE(!napi_pp_put_page(netmems[k], false)); + + kvfree(tokens); + return ret; +} +#endif + void sockopt_lock_sock(struct sock *sk) { /* When current->bpf_ctx is set, the setsockopt is called from @@ -1200,6 +1257,10 @@ int sk_setsockopt(struct sock *sk, int level, int optname, ret = -EOPNOTSUPP; return ret; } +#ifdef CONFIG_PAGE_POOL + case SO_DEVMEM_DONTNEED: + return sock_devmem_dontneed(sk, optval, optlen); +#endif } sockopt_lock_sock(sk); -- 2.44.0.396.g6e790dbe36-goog