Received: by 2002:a05:6a10:6744:0:0:0:0 with SMTP id w4csp4698136pxu; Wed, 21 Oct 2020 03:05:58 -0700 (PDT) X-Google-Smtp-Source: ABdhPJw0zXy65ojPdoUxScev8LgUc931QD7W1GB3DVdIcAKClp7vA90Z8yTCFH+Zvi7hTm6g9YIS X-Received: by 2002:a50:9e82:: with SMTP id a2mr2180385edf.117.1603274758617; Wed, 21 Oct 2020 03:05:58 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1603274758; cv=none; d=google.com; s=arc-20160816; b=Q++ISS0s8pe/6LYBnQ71Kc9m3aS8E+4vR2srEWgEJDlnoWqEEIQX7GWgqOhCkhncbp LBAHVdB4smafY3LLRNDXJM1IFrqVRKGQGMjlQ5iKi0nk5qDqjOcELP34Ytt6JK4gPB+b JorQJQ/CZlsAXrMU7coBBF1l5dh3ryTIE7acLOz2Rxrhq+urmR9AlzaXXJLl5xTByESm MCNptNZcg8t8jHSiejXHTS+Ikd1Vk0q5ynQbHJMXLRhNmAVrUgrMaSWueO4btSri55Hu i6XlkjtI/02vGA07q51V67ENFi/qh5E0a4r/uqArBoXRv2YWOnQ/RwMQZoEGKoyPgYLT wcfg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:content-transfer-encoding:mime-version :references:in-reply-to:message-id:date:subject:cc:to:from; bh=mKXeGSjs2PX4ElXU18mmsTyU03ZjqgEm7GsyImxM4Qk=; b=giQtKykZe1XkxGVkkGXczu/zCzQk3Tm8bQCQgYIey0iTDqu/AdFvjS3lUZ7D/HO7QA mjodjbU+cPHL/Eeu2YBwzpOfmRpkpu0bhuDqABm+G546MxHs1px3sOnOE9kxn1cRWKhM Bn4EKImL9hTOFgvCqsssaSSr1HbszFSO7PVmuXjJYXCLJx3S9XyOSqp5NeNChcCdvHJN /m1UGM2qhV0FCGBRZ/zMU6OC6Q95alSvJ1ZD6ZW4XtU9TlqyYfRKTf9GbiW0o4nuqRF8 Qzcrly1F+cW5CCmNEAkjXWdPwLNeKQUHA0lp8ruyifwoS4wKp3Vf0223gXZD7eXwmQYE RIMQ== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of linux-crypto-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) smtp.mailfrom=linux-crypto-owner@vger.kernel.org Return-Path: Received: from vger.kernel.org (vger.kernel.org. [23.128.96.18]) by mx.google.com with ESMTP id o21si1038686ejx.13.2020.10.21.03.05.34; Wed, 21 Oct 2020 03:05:58 -0700 (PDT) Received-SPF: pass (google.com: domain of linux-crypto-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) client-ip=23.128.96.18; Authentication-Results: mx.google.com; spf=pass (google.com: domain of linux-crypto-owner@vger.kernel.org designates 23.128.96.18 as permitted sender) smtp.mailfrom=linux-crypto-owner@vger.kernel.org Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S2438600AbgJTUkI (ORCPT + 99 others); Tue, 20 Oct 2020 16:40:08 -0400 Received: from mail-qv1-f67.google.com ([209.85.219.67]:37881 "EHLO mail-qv1-f67.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S2438591AbgJTUkG (ORCPT ); Tue, 20 Oct 2020 16:40:06 -0400 Received: by mail-qv1-f67.google.com with SMTP id t6so1621668qvz.4; Tue, 20 Oct 2020 13:40:06 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:from:to:cc:subject:date:message-id:in-reply-to :references:mime-version:content-transfer-encoding; bh=mKXeGSjs2PX4ElXU18mmsTyU03ZjqgEm7GsyImxM4Qk=; b=G0B+hpiok9ELahrTCbQds+rnuWxqyPTSBxvFqvFZPRJdts0AoAtrlDdp7UsLdu4PFo QDiWAw/GSQ8meH4SN+m+oQaxrsMeoa8MAFyOVwzyspK1AXnamiyfk/p7mMD8YLDNgaWi vtsEhpmOUzYF05Aerv7qCRa+L2Xc3g+6g30m5q3mpc+Mt4qzzakQkdx+7b6A3glj0iOV 4asqep6K9dph0XJke/PSmKRa3gyzn6kF7akCgdZjIXFzGMGSpx2LccIiZXUuCP7kcmlN B5iMR55em9yAL9r1gyHs6Y0IifcimBFzw8SqoXoQ4VzWnSkgP/bhyu9slYKWnU5ruRF8 pm1Q== X-Gm-Message-State: AOAM533z1/8CAPBsNENVMlykdXZXobXqOZj4nggmi7/9UWa47Lk1SXpX vu9ozAmSgOPcV6kcwtAsAjw= X-Received: by 2002:a0c:b251:: with SMTP id k17mr5301626qve.53.1603226405645; Tue, 20 Oct 2020 13:40:05 -0700 (PDT) Received: from rani.riverdale.lan ([2001:470:1f07:5f3::b55f]) by smtp.gmail.com with ESMTPSA id m18sm1411165qkk.102.2020.10.20.13.40.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 20 Oct 2020 13:40:04 -0700 (PDT) From: Arvind Sankar To: Herbert Xu , "David S. Miller" , "linux-crypto@vger.kernel.org" , David Laight Cc: linux-kernel@vger.kernel.org Subject: [PATCH v2 5/6] crypto: lib/sha256 - Unroll LOAD and BLEND loops Date: Tue, 20 Oct 2020 16:39:56 -0400 Message-Id: <20201020203957.3512851-6-nivedita@alum.mit.edu> X-Mailer: git-send-email 2.26.2 In-Reply-To: <20201020203957.3512851-1-nivedita@alum.mit.edu> References: <20201020203957.3512851-1-nivedita@alum.mit.edu> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Precedence: bulk List-ID: X-Mailing-List: linux-crypto@vger.kernel.org Unrolling the LOAD and BLEND loops improves performance by ~8% on x86_64 (tested on Broadwell Xeon) while not increasing code size too much. Signed-off-by: Arvind Sankar --- lib/crypto/sha256.c | 24 ++++++++++++++++++++---- 1 file changed, 20 insertions(+), 4 deletions(-) diff --git a/lib/crypto/sha256.c b/lib/crypto/sha256.c index 5efd390706c6..3a8802d5f747 100644 --- a/lib/crypto/sha256.c +++ b/lib/crypto/sha256.c @@ -68,12 +68,28 @@ static void sha256_transform(u32 *state, const u8 *input, u32 *W) int i; /* load the input */ - for (i = 0; i < 16; i++) - LOAD_OP(i, W, input); + for (i = 0; i < 16; i += 8) { + LOAD_OP(i + 0, W, input); + LOAD_OP(i + 1, W, input); + LOAD_OP(i + 2, W, input); + LOAD_OP(i + 3, W, input); + LOAD_OP(i + 4, W, input); + LOAD_OP(i + 5, W, input); + LOAD_OP(i + 6, W, input); + LOAD_OP(i + 7, W, input); + } /* now blend */ - for (i = 16; i < 64; i++) - BLEND_OP(i, W); + for (i = 16; i < 64; i += 8) { + BLEND_OP(i + 0, W); + BLEND_OP(i + 1, W); + BLEND_OP(i + 2, W); + BLEND_OP(i + 3, W); + BLEND_OP(i + 4, W); + BLEND_OP(i + 5, W); + BLEND_OP(i + 6, W); + BLEND_OP(i + 7, W); + } /* load the state into our registers */ a = state[0]; b = state[1]; c = state[2]; d = state[3]; -- 2.26.2