by Xu, Like

[permalink] [raw]

Subject: Re: [PATCH v13 02/10] KVM: x86/vmx: Make vmx_set_intercept_for_msr() non-static and expose it

Hi Sean,

On 2020/9/29 11:13, Sean Christopherson wrote:
> On Sun, Jul 26, 2020 at 11:32:21PM +0800, Like Xu wrote:
>> It's reasonable to call vmx_set_intercept_for_msr() in other vmx-specific
>> files (e.g. pmu_intel.c), so expose it without semantic changes hopefully.
> I suppose it's reasonable, but you still need to state what is actually
> going to use it.

Sure, I will add more here that
one of its usage is to pass through LBR-related msrs later.

>> Signed-off-by: Like Xu <[email protected]>
>> ---
>> arch/x86/kvm/vmx/vmx.c | 4 ++--
>> arch/x86/kvm/vmx/vmx.h | 2 ++
>> 2 files changed, 4 insertions(+), 2 deletions(-)
>>
>> diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
>> index dcde73a230c6..162c668d58f5 100644
>> --- a/arch/x86/kvm/vmx/vmx.c
>> +++ b/arch/x86/kvm/vmx/vmx.c
>> @@ -3772,8 +3772,8 @@ static __always_inline void vmx_enable_intercept_for_msr(unsigned long *msr_bitm
>> }
>> }
>>
>> -static __always_inline void vmx_set_intercept_for_msr(unsigned long *msr_bitmap,
>> - u32 msr, int type, bool value)
>> +__always_inline void vmx_set_intercept_for_msr(unsigned long *msr_bitmap,
>> + u32 msr, int type, bool value)
>> {
>> if (value)
>> vmx_enable_intercept_for_msr(msr_bitmap, msr, type);
>> diff --git a/arch/x86/kvm/vmx/vmx.h b/arch/x86/kvm/vmx/vmx.h
>> index 0d06951e607c..08c850596cfc 100644
>> --- a/arch/x86/kvm/vmx/vmx.h
>> +++ b/arch/x86/kvm/vmx/vmx.h
>> @@ -356,6 +356,8 @@ void vmx_update_host_rsp(struct vcpu_vmx *vmx, unsigned long host_rsp);
>> int vmx_find_msr_index(struct vmx_msrs *m, u32 msr);
>> int vmx_handle_memory_failure(struct kvm_vcpu *vcpu, int r,
>> struct x86_exception *e);
>> +void vmx_set_intercept_for_msr(unsigned long *msr_bitmap,
>> + u32 msr, int type, bool value);
> This completely defeats the purpose of __always_inline.
My motivation is to use vmx_set_intercept_for_msr() in pmu_intel.c,
which helps to extract pmu-specific code from vmx.c

I assume modern compilers will still make it inline even in this way,
or do you have a better solution for this?

Please let me if you have more comments on the patch series.

Thanks,
Like Xu
>
>>
>> #define POSTED_INTR_ON 0
>> #define POSTED_INTR_SN 1
>> --
>> 2.21.3
>>

2020-09-30 02:25:32

by Xu, Like

[permalink] [raw]

Subject: Re: [PATCH v13 00/10] Guest Last Branch Recording Enabling (KVM part)

Are there volunteers or maintainer to help review this patch-set ？

Just a kindly ping.

Please let me know if you need a re-based version.

Thanks,
Like Xu

On 2020/8/14 16:48, Xu, Like wrote:
> Are there no interested reviewers or users?
>
> Just a kindly ping.
>
> On 2020/7/26 23:32, Like Xu wrote:
>> Hi Paolo,
>>
>> Please review this new version for the Kernel 5.9 release, and
>> Sean may not review them as he said in the previous email
>> https://lore.kernel.org/kvm/[email protected]/
>>
>> You may cherry-pick the perf patches "3cb9d5464c1c..e1ad1ac2deb8"
>> from the branch "tip/perf/core" of scm/linux/kernel/git/tip/tip.git
>> as PeterZ said in the previous email
>> https://lore.kernel.org/kvm/[email protected]/
>>
>>
>> We may also apply the qemu-devel patch to the upstream qemu and try
>> the QEMU command lines with '-cpu host' or '-cpu host,pmu=true,lbr=true'.
>>
>> The following error will be gone forever with the patchset:
>>
>>    $ perf record -b lbr ${WORKLOAD}
>>    or $ perf record --call-graph lbr ${WORKLOAD}
>>    Error:
>>    cycles: PMU Hardware doesn't support sampling/overflow-interrupts.
>> Try 'perf stat'
>>
>> Please check more details in each commit and feel free to test.
>>
>> v12->v13 Changelog:
>> - remove perf patches since they're queued in the tip/perf/core;
>> - add a minor patch to refactor MSR_IA32_DEBUGCTLMSR set/get handler;
>> - add a minor patch to expose vmx_set_intercept_for_msr();
>> - add a minor patch to initialize perf_capabilities in the
>> intel_pmu_init();
>> - spilt the big patch to three pieces (0004-0006) for better
>> understanding and review
>> - make the LBR_FMT exposure patch as the last step to enable guest LBR;
>>
>> Previous:
>> https://lore.kernel.org/kvm/[email protected]/
>>
>>
>> ---
>>
>> The last branch recording (LBR) is a performance monitor unit (PMU)
>> feature on Intel processors that records a running trace of the most
>> recent branches taken by the processor in the LBR stack. This patch
>> series is going to enable this feature for plenty of KVM guests.
>>
>> The user space could configure whether it's enabled or not for each
>> guest via MSR_IA32_PERF_CAPABILITIES msr. As a first step, a guest
>> could only enable LBR feature if its cpu model is the same as the
>> host since the LBR feature is still one of model specific features.
>>
>> If it's enabled on the guest, the guest LBR driver would accesses the
>> LBR MSR (including IA32_DEBUGCTLMSR and records MSRs) as host does.
>> The first guest access on the LBR related MSRs is always interceptible.
>> The KVM trap would create a special LBR event (called guest LBR event)
>> which enables the callstack mode and none of hardware counter is assigned.
>> The host perf would enable and schedule this event as usual.
>>
>> Guest's first access to a LBR registers gets trapped to KVM, which
>> creates a guest LBR perf event. It's a regular LBR perf event which gets
>> the LBR facility assigned from the perf subsystem. Once that succeeds,
>> the LBR stack msrs are passed through to the guest for efficient accesses.
>> However, if another host LBR event comes in and takes over the LBR
>> facility, the LBR msrs will be made interceptible, and guest following
>> accesses to the LBR msrs will be trapped and meaningless.
>>
>> Because saving/restoring tens of LBR MSRs (e.g. 32 LBR stack entries) in
>> VMX transition brings too excessive overhead to frequent vmx transition
>> itself, the guest LBR event would help save/restore the LBR stack msrs
>> during the context switching with the help of native LBR event callstack
>> mechanism, including LBR_SELECT msr.
>>
>> If the guest no longer accesses the LBR-related MSRs within a scheduling
>> time slice and the LBR enable bit is unset, vPMU would release its guest
>> LBR event as a normal event of a unused vPMC and the pass-through
>> state of the LBR stack msrs would be canceled.
>>
>> ---
>>
>> LBR testcase:
>> echo 1 > /proc/sys/kernel/watchdog
>> echo 25 > /proc/sys/kernel/perf_cpu_time_max_percent
>> echo 5000 > /proc/sys/kernel/perf_event_max_sample_rate
>> echo 0 > /proc/sys/kernel/perf_cpu_time_max_percent
>> ./perf record -b ./br_instr a
>>
>> - Perf report on the host:
>> Samples: 72K of event 'cycles', Event count (approx.): 72512
>> Overhead Command   Source Shared Object           Source
>> Symbol                           Target Symbol
>> Basic Block Cycles
>>    12.12% br_instr br_instr                       [.]
>> cmp_end                             [.]
>> lfsr_cond                           1
>>    11.05% br_instr br_instr                       [.]
>> lfsr_cond                           [.]
>> cmp_end                             5
>>     8.81% br_instr br_instr                       [.]
>> lfsr_cond                           [.]
>> cmp_end                             4
>>     5.04% br_instr br_instr                       [.]
>> cmp_end                             [.]
>> lfsr_cond                           20
>>     4.92% br_instr br_instr                       [.]
>> lfsr_cond                           [.]
>> cmp_end                             6
>>     4.88% br_instr br_instr                       [.]
>> cmp_end                             [.]
>> lfsr_cond                           6
>>     4.58% br_instr br_instr                       [.]
>> cmp_end                             [.]
>> lfsr_cond                           5
>>
>> - Perf report on the guest:
>> Samples: 92K of event 'cycles', Event count (approx.): 92544
>> Overhead Command   Source Shared Object Source
>> Symbol                                   Target
>> Symbol                                   Basic Block Cycles
>>    12.03% br_instr br_instr              [.]
>> cmp_end                                     [.]
>> lfsr_cond                                   1
>>    11.09% br_instr br_instr              [.]
>> lfsr_cond                                   [.]
>> cmp_end                                     5
>>     8.57% br_instr br_instr              [.]
>> lfsr_cond                                   [.]
>> cmp_end                                     4
>>     5.08% br_instr br_instr              [.]
>> lfsr_cond                                   [.]
>> cmp_end                                     6
>>     5.06% br_instr br_instr              [.]
>> cmp_end                                     [.]
>> lfsr_cond                                   20
>>     4.87% br_instr br_instr              [.]
>> cmp_end                                     [.]
>> lfsr_cond                                   6
>>     4.70% br_instr br_instr              [.]
>> cmp_end                                     [.]
>> lfsr_cond                                   5
>>
>> Conclusion: the profiling results on the guest are similar to that on
>> the host.
>>
>> Like Xu (10):
>>    KVM: x86: Move common set/get handler of MSR_IA32_DEBUGCTLMSR to VMX
>>    KVM: x86/vmx: Make vmx_set_intercept_for_msr() non-static and expose it
>>    KVM: vmx/pmu: Initialize vcpu perf_capabilities once in intel_pmu_init()
>>    KVM: vmx/pmu: Clear PMU_CAP_LBR_FMT when guest LBR is disabled
>>    KVM: vmx/pmu: Create a guest LBR event when vcpu sets DEBUGCTLMSR_LBR
>>    KVM: vmx/pmu: Pass-through LBR msrs to when the guest LBR event is
>> ACTIVE
>>    KVM: vmx/pmu: Reduce the overhead of LBR pass-through or cancellation
>>    KVM: vmx/pmu: Emulate legacy freezing LBRs on virtual PMI
>>    KVM: vmx/pmu: Expose LBR_FMT in the MSR_IA32_PERF_CAPABILITIES
>>    KVM: vmx/pmu: Release guest LBR event via lazy release mechanism
>>
>> arch/x86/kvm/pmu.c              | 12 +-
>> arch/x86/kvm/pmu.h              |   5 +
>> arch/x86/kvm/vmx/capabilities.h | 22 ++-
>> arch/x86/kvm/vmx/pmu_intel.c    | 296 +++++++++++++++++++++++++++++++-
>> arch/x86/kvm/vmx/vmx.c          | 44 ++++-
>> arch/x86/kvm/vmx/vmx.h          | 28 +++
>> arch/x86/kvm/x86.c              | 15 +-
>> 7 files changed, 395 insertions(+), 27 deletions(-)
>>
>