Received: by 2002:a25:c593:0:0:0:0:0 with SMTP id v141csp1538012ybe; Wed, 11 Sep 2019 17:00:37 -0700 (PDT) X-Google-Smtp-Source: APXvYqy2nhPbHdcjIdjuRYwckIw+HohBYV/BtbpitEeWewqdiprqiIepc3WnbWLi0/E4VQ0dK8gd X-Received: by 2002:aa7:ce99:: with SMTP id y25mr2591175edv.145.1568246437769; Wed, 11 Sep 2019 17:00:37 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1568246437; cv=none; d=google.com; s=arc-20160816; b=V8Tmnhl/2MqCsItiDVrxumKYqK+9Vj+vctjxZQF3TNTB5BfxveRfDyAcJeHo29tGT7 LhfAbC6hyFsbeweKvKXTaVR0Vv4nP92iGKAaedj4yynozbObQC6Y+TLCuMYxjkJqdWR8 tpzeHUP3dBhOsp9c7w607THDJq1OPI6sD6hnjk+VWSyZaBaca4iY6132DMtNNZ8mKIE5 CuaOIRgIr6wBFW/rxQtfM1PBzB68wCAqopp0hOvnBKWTo1+y52M5G3c80PR5TKTpd+kG bCfWrmvX3MP3VC7EYm6qUL+PGMSNMN0ol+ujeGk+PzKn3grVCVQ7jya/QinG3dfhypyp AypQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:sender:content-language :content-transfer-encoding:in-reply-to:mime-version:user-agent:date :message-id:organization:from:references:cc:to:subject; bh=impvmOEbjCN9y7I8NH2OhVk1+5WlZGKThLqnE34IH6k=; b=TGsRyLpjnv2g/amvgzwIjL7EzCORHrF3EtoIlJn+ZuqZpNfH/NVQwpdeFinyPTUbZv CzxMwmTsad8OV14jAipa1aWBW1YSxpU63cJKeoQSUNk0F5Cn284aJ459ErxbM/BW8EDK lPmSAfrjR5FiXvYoilisK4Z+b0jn5UAFP9kr4w+reBCf9ohKekHKsLmkAgUdz+yMcLOw yVflMmhw+aKFqAJWXwwuaNcPXfPRQAOEGTHbPbhd3C4PDp8SwjRJfHhs2fa4sNvBU1Qt 4bqqEKnbsn5dXdZdY9yisT9rQCRM3JowPPJ9S4DgmOR2ZSB5dPjGdVl+9u+O9juCp6b/ 1SMw== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=redhat.com Return-Path: Received: from vger.kernel.org (vger.kernel.org. [209.132.180.67]) by mx.google.com with ESMTP id n13si13053416edq.98.2019.09.11.16.59.58; Wed, 11 Sep 2019 17:00:37 -0700 (PDT) Received-SPF: pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) client-ip=209.132.180.67; Authentication-Results: mx.google.com; spf=pass (google.com: best guess record for domain of linux-kernel-owner@vger.kernel.org designates 209.132.180.67 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=redhat.com Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1729656AbfIKUyF (ORCPT + 99 others); Wed, 11 Sep 2019 16:54:05 -0400 Received: from mx1.redhat.com ([209.132.183.28]:33966 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1728412AbfIKUyF (ORCPT ); Wed, 11 Sep 2019 16:54:05 -0400 Received: from smtp.corp.redhat.com (int-mx04.intmail.prod.int.phx2.redhat.com [10.5.11.14]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by mx1.redhat.com (Postfix) with ESMTPS id 1734D3091785; Wed, 11 Sep 2019 20:54:04 +0000 (UTC) Received: from llong.remote.csb (ovpn-121-77.rdu2.redhat.com [10.10.121.77]) by smtp.corp.redhat.com (Postfix) with ESMTP id 46B845D9E2; Wed, 11 Sep 2019 20:54:01 +0000 (UTC) Subject: Re: [PATCH 5/5] hugetlbfs: Limit wait time when trying to share huge PMD To: Qian Cai Cc: Peter Zijlstra , Ingo Molnar , Will Deacon , Alexander Viro , Mike Kravetz , linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, Davidlohr Bueso References: <20190911150537.19527-1-longman@redhat.com> <20190911150537.19527-6-longman@redhat.com> <1a8e6c0a-6ba6-d71f-974e-f8a9c623c25b@redhat.com> <70714929-2CE3-42F4-BD31-427077C9E24E@lca.pw> From: Waiman Long Organization: Red Hat Message-ID: <211b144f-0e86-d891-e1ec-9879ceb53e36@redhat.com> Date: Wed, 11 Sep 2019 21:54:00 +0100 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Thunderbird/60.7.2 MIME-Version: 1.0 In-Reply-To: <70714929-2CE3-42F4-BD31-427077C9E24E@lca.pw> Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit Content-Language: en-US X-Scanned-By: MIMEDefang 2.79 on 10.5.11.14 X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.41]); Wed, 11 Sep 2019 20:54:04 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 9/11/19 8:42 PM, Qian Cai wrote: > >> On Sep 11, 2019, at 12:34 PM, Waiman Long wrote: >> >> On 9/11/19 5:01 PM, Qian Cai wrote: >>>> On Sep 11, 2019, at 11:05 AM, Waiman Long wrote: >>>> >>>> When allocating a large amount of static hugepages (~500-1500GB) on a >>>> system with large number of CPUs (4, 8 or even 16 sockets), performance >>>> degradation (random multi-second delays) was observed when thousands >>>> of processes are trying to fault in the data into the huge pages. The >>>> likelihood of the delay increases with the number of sockets and hence >>>> the CPUs a system has. This only happens in the initial setup phase >>>> and will be gone after all the necessary data are faulted in. >>>> >>>> These random delays, however, are deemed unacceptable. The cause of >>>> that delay is the long wait time in acquiring the mmap_sem when trying >>>> to share the huge PMDs. >>>> >>>> To remove the unacceptable delays, we have to limit the amount of wait >>>> time on the mmap_sem. So the new down_write_timedlock() function is >>>> used to acquire the write lock on the mmap_sem with a timeout value of >>>> 10ms which should not cause a perceivable delay. If timeout happens, >>>> the task will abandon its effort to share the PMD and allocate its own >>>> copy instead. >>>> >>>> When too many timeouts happens (threshold currently set at 256), the >>>> system may be too large for PMD sharing to be useful without undue delay. >>>> So the sharing will be disabled in this case. >>>> >>>> Signed-off-by: Waiman Long >>>> --- >>>> include/linux/fs.h | 7 +++++++ >>>> mm/hugetlb.c | 24 +++++++++++++++++++++--- >>>> 2 files changed, 28 insertions(+), 3 deletions(-) >>>> >>>> diff --git a/include/linux/fs.h b/include/linux/fs.h >>>> index 997a530ff4e9..e9d3ad465a6b 100644 >>>> --- a/include/linux/fs.h >>>> +++ b/include/linux/fs.h >>>> @@ -40,6 +40,7 @@ >>>> #include >>>> #include >>>> #include >>>> +#include >>>> >>>> #include >>>> #include >>>> @@ -519,6 +520,12 @@ static inline void i_mmap_lock_write(struct address_space *mapping) >>>> down_write(&mapping->i_mmap_rwsem); >>>> } >>>> >>>> +static inline bool i_mmap_timedlock_write(struct address_space *mapping, >>>> + ktime_t timeout) >>>> +{ >>>> + return down_write_timedlock(&mapping->i_mmap_rwsem, timeout); >>>> +} >>>> + >>>> static inline void i_mmap_unlock_write(struct address_space *mapping) >>>> { >>>> up_write(&mapping->i_mmap_rwsem); >>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c >>>> index 6d7296dd11b8..445af661ae29 100644 >>>> --- a/mm/hugetlb.c >>>> +++ b/mm/hugetlb.c >>>> @@ -4750,6 +4750,8 @@ void adjust_range_if_pmd_sharing_possible(struct vm_area_struct *vma, >>>> } >>>> } >>>> >>>> +#define PMD_SHARE_DISABLE_THRESHOLD (1 << 8) >>>> + >>>> /* >>>> * Search for a shareable pmd page for hugetlb. In any case calls pmd_alloc() >>>> * and returns the corresponding pte. While this is not necessary for the >>>> @@ -4770,11 +4772,24 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud) >>>> pte_t *spte = NULL; >>>> pte_t *pte; >>>> spinlock_t *ptl; >>>> + static atomic_t timeout_cnt; >>>> >>>> - if (!vma_shareable(vma, addr)) >>>> - return (pte_t *)pmd_alloc(mm, pud, addr); >>>> + /* >>>> + * Don't share if it is not sharable or locking attempt timed out >>>> + * after 10ms. After 256 timeouts, PMD sharing will be permanently >>>> + * disabled as it is just too slow. >>> It looks like this kind of policy interacts with kernel debug options like KASAN (which is going to slow the system down >>> anyway) could introduce tricky issues due to different timings on a debug kernel. >> With respect to lockdep, down_write_timedlock() works like a trylock. So >> a lot of checking will be skipped. Also the lockdep code won't be run >> until the lock is acquired. So its execution time has no effect on the >> timeout. > No only lockdep, but also things like KASAN, debug_pagealloc, page_poison, kmemleak, debug > objects etc that all going to slow down things in huge_pmd_share(), and make it tricky to get a > right timeout value for those debug kernels without changing the previous behavior. Right, I understand that. I will move to use a sysctl parameters for the timeout and then set its default value to either 10ms or 20ms if some debug options are detected. Usually the slower than should not be more than 2X. Cheers, Longman