Received: by 2002:a05:6358:9144:b0:117:f937:c515 with SMTP id r4csp7206198rwr; Tue, 2 May 2023 10:58:51 -0700 (PDT) X-Google-Smtp-Source: ACHHUZ7zThG5abeifgY4oHe7dvOnniq7hEn5Kzj5AmzRLwboMcIDLr7PdIJ5svl0CFuYRwPudNfK X-Received: by 2002:a05:6a21:9991:b0:ef:7d7b:433b with SMTP id ve17-20020a056a21999100b000ef7d7b433bmr22615246pzb.41.1683050330999; Tue, 02 May 2023 10:58:50 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1683050330; cv=none; d=google.com; s=arc-20160816; b=LDagRKpaZMsFCUspk+6QHIk1+q1xwwz6J5tFw0wjOtVP2KDRRUf2ZHNo5PmJn5IhbD C/0o3B0BBF0BiyMKXgfmf0u9OSd22tjzEda7WhkTRJCoV70PwFrd9KNx6BugLYgVJsnj nBCeBmMIj8OhagkekluXmIgmi+5l30yCCQemeii0j52sHwgJZtowtEkptN9Qm3XI8Emq zA9V9YI0Cfj6ggzhv+6Yr/GtQwV6VEcC23gX8BIK3XnWphxM3Ic9YcduYmZvob6pq5y5 A1suVmwJMCbs/S5kEa9+g9zA4BTiQOUIoP8JcR1HhFdJeefa189Ms9MfmZP3GFauGubi YSbQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:in-reply-to:content-disposition:mime-version :references:message-id:subject:cc:to:from:date:dkim-signature; bh=/465ioAn95NNEKxI/y1OxBfAvlwiRZKFhD46N2OZRqk=; b=lfl8/mVRBiZAgG4RJPv5YSFt3oKid5xUbAdX2n2U+ygAUPBKcklvc6XlsMQzkfL+o3 a4/CBBeUrfZxeyunJdovidV8CX8u+nMNyCuMeS6uT9eL9Ox2chinslULK2YblM44W6m2 s964j9YFqRu+v1BuikwMm3VBTXfXXiLdfxnQDfYORr3qe2QJxOeXFuUYzl9tnlKAb6KL ubl/yUgpITSbAxx+nsyDW69od14Syo5KfRvutnCg6ttjvUp8AMpo5T3Orqb+K20w6UAc 9zc4F05oDsE8tvsWCsszKu7O0MnxjHVZLDubpDZ73GzEZG/A0Oo9Zha7Rm9QX1nbwp9D ApLA== ARC-Authentication-Results: i=1; mx.google.com; dkim=pass header.i=@gmail.com header.s=20221208 header.b=MDLyOim5; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=pass (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Return-Path: Received: from out1.vger.email (out1.vger.email. [2620:137:e000::1:20]) by mx.google.com with ESMTP id k191-20020a6384c8000000b00513b4eb72a5si32256310pgd.731.2023.05.02.10.58.38; Tue, 02 May 2023 10:58:50 -0700 (PDT) Received-SPF: pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) client-ip=2620:137:e000::1:20; Authentication-Results: mx.google.com; dkim=pass header.i=@gmail.com header.s=20221208 header.b=MDLyOim5; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=pass (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234420AbjEBRpw (ORCPT + 99 others); Tue, 2 May 2023 13:45:52 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:40166 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229575AbjEBRpu (ORCPT ); Tue, 2 May 2023 13:45:50 -0400 Received: from mail-wr1-x433.google.com (mail-wr1-x433.google.com [IPv6:2a00:1450:4864:20::433]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id E4CE3C3; Tue, 2 May 2023 10:45:47 -0700 (PDT) Received: by mail-wr1-x433.google.com with SMTP id ffacd0b85a97d-3062c1e7df8so1908144f8f.1; Tue, 02 May 2023 10:45:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20221208; t=1683049546; x=1685641546; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=/465ioAn95NNEKxI/y1OxBfAvlwiRZKFhD46N2OZRqk=; b=MDLyOim5BMRn6hfXjaVej53KW+5dmIM3Kl9aUN5vg5pGtYXJB3mydYvJYPbHAvr5W0 L8Nl8T5KFv+aMHZZ1EUjV0xn5gHIQQLJ8t/0RFiB3xgixHqpsIBsnCtLoPZ4M0IdEkae iFnNOgkq4akftSubl/G37wbbq8Ryyeyqjalcc/nswPnpN2ygGGeD/MNZpVYD5dQwKmn3 Zw2waXt5myN4W17JESf1xHHsPzjuYjD9uQ8LE6SMACY0gesoDxJvnPXk89JVk2qVjwmn sJ6SWbpbNvuodR5JVQGnq1AfcxkDsORytjYHKofxm8N9wP29zJW0uKO4EHE9s6x0CyzG GzNQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20221208; t=1683049546; x=1685641546; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=/465ioAn95NNEKxI/y1OxBfAvlwiRZKFhD46N2OZRqk=; b=ceQFAQ8hgXMJ8EMjf9oPuGecP19B+50yWH/j8g7UmuJNrow7YFoecdkIUWpDKJ05NM 9nK9PTBulx+xKxBqn5OcwnNPrRmagiSBwDydw9dzsiCxj9rg05h+78thPbNvqEHxZxSQ xbky19i43AxaZrn7BnuBdhv3LfPlRoc2QDORyDwIsx0V1lvGbGZBn/w7P2sX+YKOMQdr EkCoYY1nfAEx1jsaQD6tFwROIEezzK3iJ2HRH8Ir4EeI/qSRZ8aIEKH2hx6tTn2KcoRO ZLJvvj4qelTNpUgvAF9XKhSlZp1YG8pXUSfFIWhbCd/A7KzXJNYvQaTjiFBg82pMz9Zy DRJA== X-Gm-Message-State: AC+VfDwAfFNhZyaqFScTLnB/qVu01yO7l5uOQqKLa8w9buO1HIG+QzBZ 8p1kVWsyb6M53faYa24IifU= X-Received: by 2002:adf:ce05:0:b0:306:34f6:de85 with SMTP id p5-20020adfce05000000b0030634f6de85mr2709622wrn.58.1683049546176; Tue, 02 May 2023 10:45:46 -0700 (PDT) Received: from localhost (host86-156-84-164.range86-156.btcentralplus.com. [86.156.84.164]) by smtp.gmail.com with ESMTPSA id f15-20020a7bcd0f000000b003f182cc55c4sm36031959wmj.12.2023.05.02.10.45.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 02 May 2023 10:45:45 -0700 (PDT) Date: Tue, 2 May 2023 18:45:44 +0100 From: Lorenzo Stoakes To: David Hildenbrand Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, Andrew Morton , Jason Gunthorpe , Jens Axboe , Matthew Wilcox , Dennis Dalessandro , Leon Romanovsky , Christian Benvenuti , Nelson Escobar , Bernard Metzler , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Mark Rutland , Alexander Shishkin , Jiri Olsa , Namhyung Kim , Ian Rogers , Adrian Hunter , Bjorn Topel , Magnus Karlsson , Maciej Fijalkowski , Jonathan Lemon , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Christian Brauner , Richard Cochran , Alexei Starovoitov , Daniel Borkmann , Jesper Dangaard Brouer , John Fastabend , linux-fsdevel@vger.kernel.org, linux-perf-users@vger.kernel.org, netdev@vger.kernel.org, bpf@vger.kernel.org, Oleg Nesterov , Jason Gunthorpe , John Hubbard , Jan Kara , "Kirill A . Shutemov" , Pavel Begunkov , Mika Penttila , Dave Chinner , Theodore Ts'o , Peter Xu , Matthew Rosato , "Paul E . McKenney" , Christian Borntraeger , Mike Rapoport Subject: Re: [PATCH v7 3/3] mm/gup: disallow FOLL_LONGTERM GUP-fast writing to file-backed mappings Message-ID: <88fcd103-7302-4838-a730-f7e0f189cfe7@lucifer.local> References: <1691115d-dba4-636b-d736-6a20359a67c3@redhat.com> <392debc7-2de8-440e-8b26-20f2d42cdf8d@lucifer.local> <6f17af6b-0925-12bd-5041-14462dab2768@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <6f17af6b-0925-12bd-5041-14462dab2768@redhat.com> X-Spam-Status: No, score=-2.1 required=5.0 tests=BAYES_00,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,DKIM_VALID_EF,FREEMAIL_FROM, RCVD_IN_DNSWL_NONE,SPF_HELO_NONE,SPF_PASS,T_SCC_BODY_TEXT_LINE autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on lindbergh.monkeyblade.net Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Tue, May 02, 2023 at 07:38:27PM +0200, David Hildenbrand wrote: > On 02.05.23 19:31, Lorenzo Stoakes wrote: > > On Tue, May 02, 2023 at 07:13:49PM +0200, David Hildenbrand wrote: > > > [...] > > > > > > > +{ > > > > + struct address_space *mapping; > > > > + > > > > + /* > > > > + * GUP-fast disables IRQs - this prevents IPIs from causing page tables > > > > + * to disappear from under us, as well as preventing RCU grace periods > > > > + * from making progress (i.e. implying rcu_read_lock()). > > > > + * > > > > + * This means we can rely on the folio remaining stable for all > > > > + * architectures, both those that set CONFIG_MMU_GATHER_RCU_TABLE_FREE > > > > + * and those that do not. > > > > + * > > > > + * We get the added benefit that given inodes, and thus address_space, > > > > + * objects are RCU freed, we can rely on the mapping remaining stable > > > > + * here with no risk of a truncation or similar race. > > > > + */ > > > > + lockdep_assert_irqs_disabled(); > > > > + > > > > + /* > > > > + * If no mapping can be found, this implies an anonymous or otherwise > > > > + * non-file backed folio so in this instance we permit the pin. > > > > + * > > > > + * shmem and hugetlb mappings do not require dirty-tracking so we > > > > + * explicitly whitelist these. > > > > + * > > > > + * Other non dirty-tracked folios will be picked up on the slow path. > > > > + */ > > > > + mapping = folio_mapping(folio); > > > > + return !mapping || shmem_mapping(mapping) || folio_test_hugetlb(folio); > > > > > > "Folios in the swap cache return the swap mapping" -- you might disallow > > > pinning anonymous pages that are in the swap cache. > > > > > > I recall that there are corner cases where we can end up with an anon page > > > that's mapped writable but still in the swap cache ... so you'd fallback to > > > the GUP slow path (acceptable for these corner cases, I guess), however > > > especially the comment is a bit misleading then. > > > > How could that happen? > > > > > > > > So I'd suggest not dropping the folio_test_anon() check, or open-coding it > > > ... which will make this piece of code most certainly easier to get when > > > staring at folio_mapping(). Or to spell it out in the comment (usually I > > > prefer code over comments). > > > > I literally made this change based on your suggestion :) but perhaps I > > misinterpreted what you meant. > > > > I do spell it out in the comment that the page can be anonymous, But perhaps > > explicitly checking the mapping flags is the way to go. > > > > > > > > > +} > > > > + > > > > /** > > > > * try_grab_folio() - Attempt to get or pin a folio. > > > > * @page: pointer to page to be grabbed > > > > @@ -123,6 +170,8 @@ static inline struct folio *try_get_folio(struct page *page, int refs) > > > > */ > > > > struct folio *try_grab_folio(struct page *page, int refs, unsigned int flags) > > > > { > > > > + bool is_longterm = flags & FOLL_LONGTERM; > > > > + > > > > if (unlikely(!(flags & FOLL_PCI_P2PDMA) && is_pci_p2pdma_page(page))) > > > > return NULL; > > > > @@ -136,8 +185,7 @@ struct folio *try_grab_folio(struct page *page, int refs, unsigned int flags) > > > > * right zone, so fail and let the caller fall back to the slow > > > > * path. > > > > */ > > > > - if (unlikely((flags & FOLL_LONGTERM) && > > > > - !is_longterm_pinnable_page(page))) > > > > + if (unlikely(is_longterm && !is_longterm_pinnable_page(page))) > > > > return NULL; > > > > /* > > > > @@ -148,6 +196,16 @@ struct folio *try_grab_folio(struct page *page, int refs, unsigned int flags) > > > > if (!folio) > > > > return NULL; > > > > + /* > > > > + * Can this folio be safely pinned? We need to perform this > > > > + * check after the folio is stabilised. > > > > + */ > > > > + if ((flags & FOLL_WRITE) && is_longterm && > > > > + !folio_longterm_write_pin_allowed(folio)) { > > > > + folio_put_refs(folio, refs); > > > > + return NULL; > > > > + } > > > > > > So we perform this change before validating whether the PTE changed. > > > > > > Hmm, naturally, I would have done it afterwards. > > > > > > IIRC, without IPI syncs during TLB flush (i.e., > > > CONFIG_MMU_GATHER_RCU_TABLE_FREE), there is the possibility that > > > (1) We lookup the pte > > > (2) The page was unmapped and free > > > (3) The page gets reallocated and used > > > (4) We pin the page > > > (5) We dereference page->mapping > > > > But we have an implied RCU lock from disabled IRQs right? Unless that CONFIG > > option does something odd (I've not really dug into its brehaviour). It feels > > like that would break GUP-fast as a whole. > > > > > > > > If we then de-reference page->mapping that gets used by whoever allocated it > > > for something completely different (not a pointer to something reasonable), > > > I wonder if we might be in trouble. > > > > > > Checking first, whether the PTE changed makes sure that what we pinned and > > > what we're looking at is what we expected. > > > > > > ... I can spot that the page_is_secretmem() check is also done before that. > > > But it at least makes sure that it's still an LRU page before staring at the > > > mapping (making it a little safer?). > > > > As do we :) > > > > We also via try_get_folio() check to ensure that we aren't subject to a split. > > > > > > > > BUT, I keep messing up this part of the story. Maybe it all works as > > > expected because we will be synchronizing RCU somehow before actually > > > freeing the page in the !IPI case. ... but I think that's only true for page > > > tables with CONFIG_MMU_GATHER_RCU_TABLE_FREE. > > > > My understanding based on what Peter said is that the IRQs being disabled should > > prevent anything bad from happening here. > > > ... only if we verified that the PTE didn't change IIUC. IRQs disabled only > protect you from the mapping getting freed and reused (because mappings are > freed via RCU IIUC). > > But as far as I can tell, it doesn't protect you from the page itself > getting freed and reused, and whoever freed the page uses page->mapping to > store something completely different. Ack, and we'd not have mapping->inode to save us in an anon case either. I'd rather be as cautious as we can possibly be, so let's move this to after the 'PTE is the same' check then, will fix on respin. > > But, again, it's all complicated and confusing to me. > It's just a fiddly, complicated, delicate area I feel :) hence why I endeavour to take on board the community's views on this series to ensure we end up with the best possible implementation. > > page_is_secretmem() also doesn't use a READ_ONCE() ... Perhaps one for a follow up patch... > > -- > Thanks, > > David / dhildenb >