cloud-hypervisor

mirror of https://github.com/cloud-hypervisor/cloud-hypervisor.git synced 2024-12-26 23:55:18 +00:00

Author	SHA1	Message	Date
Samuel Ortiz	fadeb98c67	cargo: Bulk update Includes updates for ssh2, cc, syn, tinyvec, backtrace micro-http and libssh2. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-11-23 12:25:31 +01:00
Rob Bradford	0fec326582	hypervisor, vmm: Remove shared ownership of VmmOps This interface is used by the vCPU thread to delegate responsibility for handling MMIO/PIO operations and to support different approaches than a VM exit. During profiling I found that we were spending 13.75% of the boot CPU uage acquiring access to the object holding the VmmOps via ArcSwap::load_full() 13.75% 6.02% vcpu0 cloud-hypervisor [.] arc_swap::ArcSwapAny<T,S>::load_full \| ---arc_swap::ArcSwapAny<T,S>::load_full \| --13.43%--<hypervisor::kvm::KvmVcpu as hypervisor::cpu::Vcpu>::run std::sys_common::backtrace::__rust_begin_short_backtrace core::ops::function::FnOnce::call_once{{vtable-shim}} std::sys::unix:🧵:Thread:🆕:thread_start However since the object implementing VmmOps does not need to be mutable and it is only used from the vCPU side we can change the ownership to being a simple Arc<> that is passed in when calling create_vcpu(). This completely removes the above CPU usage from subsequent profiles. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-19 00:16:02 +01:00
Rob Bradford	c622278030	config, device_manager: Add support for disabling io_uring for testing Add config parameter to --disk called "_disable_io_uring" (the underscore prefix indicating it is not for public consumpion.) Use this option to disable io_uring if it would otherwise be used. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-18 11:47:54 +01:00
Rob Bradford	3ac9b6c404	vmm: Implement live migration Now the VM is paused/resumed by the migration process itself. 0. The guest configuration is sent to the destination 1. Dirty page log tracking is started by start_memory_dirty_log() 2. All guest memory is sent to the destination 3. Up to 5 attempts are made to send the dirty guest memory to the destination... 4. ...before the VM is paused 5. One last set of dirty pages is sent to the destination 6. The guest is snapshotted and sent to the destination 7. When the migration is completed the destination unpauses the received VM. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Rob Bradford	b34703d29f	vmm: vm: Add dirty log related passthrough methods This allows code running in the VMM to access the VM's MemoryManager's functionality for managing the dirty log including resetting it but also generating a table. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Rob Bradford	cf6763dfdb	vmm: migration: Add missing response check A read and check of the response was missing from when sending the memory to the destination. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Rob Bradford	11a69450ba	vm-migration, vmm: Send configuration in separate step Prior to sending the memory the full state is not needed only the configuration. This is sufficient to create the appropriate structures in the guest and have the memory allocations ready for filling. Update the protocol documentation to add a separate config step and move the state to after the memory is transferred. As the VM is created in a separate step to restoring it the requires a slightly different constructor as well as saving the VM object for the subsequent commands. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Rob Bradford	c62e409827	memory_manager: Generate a MemoryRangeTable for dirty ranges In order to do this we must extend the MemoryManager API to add the ability to specify the tracking of the dirty pages when creating the userspace mappings and also keep track of the userspace mappings that have been created for RAM regions. Currently the dirty pages are collected into ranges based on a block level of 64 pages. The algorithm could be tweaked to create smaller ranges but for now if any page in the block of 64 is dirty the whole block is added to the range. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Rob Bradford	8baa244ec1	hypervisor: Add control for dirty page logging When creating a userspace mapping provide a control for enabling the logging of dirty pages. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-17 16:57:11 +00:00
Anatol Belski	b399287430	memory_manager: Make addressable space size 64k aligned While the addressable space size reduction of 4k in necessary due to the Linux bug, the 64k alignment of the addressable space size is required by Windows. This patch satisfies both. Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>	2020-11-16 16:39:11 +00:00
Rob Bradford	ca60adda70	vmm: Add support for sending and receiving migration if VM is paused This is tested by: Source VMM: target/debug/cloud-hypervisor --kernel ~/src/linux/vmlinux \ --pmem file=~/workloads/focal.raw --cpus boot=1 \ --memory size=2048M \ --cmdline"root=/dev/pmem0p1 console=ttyS0" --serial tty --console off \ --api-socket=/tmp/api1 -v Destination VMM: target/debug/cloud-hypervisor --api-socket=/tmp/api2 -v And the following commands: target/debug/ch-remote --api-socket=/tmp/api1 pause target/debug/ch-remote --api-socket=/tmp/api2 receive-migration unix:/tmp/foo & target/debug/ch-remote --api-socket=/tmp/api1 send-migration unix:/tmp/foo target/debug/ch-remote --api-socket=/tmp/api2 resume The VM is then responsive on the destination VMM. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-11 11:07:24 +01:00
Rob Bradford	dfe2dadb3e	vmm: memory_manager: Make the snapshot source directory an Option This allows the code to be reused when creating the VM from a snapshot when doing VM migration. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-11 11:07:24 +01:00
Rob Bradford	7ac764518c	vmm: api: Implement API support for migration Add API entry points with stub implementation for sending and receiving a VM from one VMM to another. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-11 11:07:24 +01:00
Julio Montes	270631922d	vmm: openapi: remove `omitempty` json tag Due to a known limitation in OpenAPITools/openapi-generator tool, it's impossible to send go zero types, like false and 0 to cloud-hypervisor because `omitempty` is added if a field is not required. Set cache_size, dax, num_queues and queue_size as required to remove `omitempty` from the json tag. fixes #1961 Signed-off-by: Julio Montes <julio.montes@intel.com>	2020-11-10 19:09:17 +01:00
Rob Bradford	7b77f1ef90	vmm: Remove self-spawning functionality for vhost-user-{net,block} This also removes the need to lookup up the "exe" symlink for finding the VMM executable path. Fixes: #1925 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-09 00:16:15 +01:00
Rob Bradford	0005d11e32	vmm: config: Require a socket when using vhost-user With self-spawning being removed both parameters are now required. Fixes: #1925 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-11-09 00:16:15 +01:00
Michael Zhao	093a581ee1	vmm: Implement VM rebooting on AArch64 The logic to handle AArch64 system event was: SHUTDOWN and RESET were all treated as RESET. Now we handle them differently: - RESET event will trigger Vmm::vm_reboot(), - SHUTDOWN event will trigger Vmm::vm_shutdown(). Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-10-30 17:14:44 +00:00
Michael Zhao	69394c9c35	vmm: Handle hypervisor VCPU run result from Vcpu to VcpuManager Now Vcpu::run() returns a boolean value to VcpuManager, indicating whether the VM is going to reboot (false) or just continue (true). Moving the handling of hypervisor VCPU run result from Vcpu to VcpuManager gives us the flexibility to handle more scenarios like shutting down on AArch64. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-10-30 17:14:44 +00:00
Rob Bradford	cb88ceeae8	vmm: memory_manager: Move the restoration of guest memory later Rather than filling the guest memory from a file at the point of the the guest memory region being created instead fill from the file later. This simplifies the region creation code but also adds flexibility for sourcing the guest memory from a source other than an on disk file. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-30 12:31:47 +01:00
Rob Bradford	21db6f53c8	vmm: memory_manager: Write all guest region to disk As a mirror of `bdbea19e23` which ensured that GuestMemoryMmap::read_exact_from() was used to read all the file to the region ensure that all the guest memory region is written to disk. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-27 12:11:31 -07:00
Rob Bradford	dfd21cbfc5	vmm: Use thiserror/anyhow for vmm::Error This gives a nicer user experience and this error can now be used as the source for other errors based off this. See: #1910 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-27 13:27:23 +00:00
Sebastien Boeuf	7e127df415	vmm: memory_manager: Replace 'ext_region' by 'saved_region' Any occurrence of of a variable containing `ext_region` is replaced with the less confusing name `saved_region`. The point is to clearly identify the memory regions that might have been saved during a snapshot, while the `ext` standing for `external` was pretty unclear. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-23 21:59:52 +02:00
Sebastien Boeuf	c0e8e5b53f	vmm: memory_manager: Replace 'backing_file' variable names In the context of saving the memory regions content through snapshot, using the term "backing file" brings confusion with the actual backing file that might back the memory mapping. To avoid such conflicting naming, the 'backing_file' field from the MemoryRegion structure gets replaced with 'content', as this is designating the potential file containing the memory region data. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-23 21:59:52 +02:00
Rob Bradford	bdbea19e23	vmm: memory_manager: Completely fill guest ram from snapshot Use GuestRegionMmap::read_exact_from() to ensure that all of the file is read into the guest. This addresses an issue where GuestRegionMmap::read_from() was only copying the first 2GiB of the memory and so lead to snapshot-restore was failing when the guest RAM was 2GiB or greater. This change also propagates any error from the copying upwards. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-23 17:56:19 +01:00
Rob Bradford	a60b437f89	vmm: memory_manager: Always copy anonymous RAM regions from disk When restoring if a region of RAM is backed by anonymous memory i.e from memfd_create() then copy the contents of the ram from the file that has been saved to disk. Previously the code would map the memory from that file into the guest using a MAP_PRIVATE mapping. This has the effect of minimising the restore time but provides an issue where the restored VM does not have the same structure as the snapshotted VM, in particular memory is backed by files in the restored VM that were anonymously backed in the original. This creates two problems: * The snapshot data is mapped from files for the pages of the guest which prevents the storage from being reclaimed. * When snapshotting again the guest memory will not be correctly saved as it will have looked like it was backed by a file so it will not be written to disk but as it is a MAP_PRIVATE mapping the changes will never be written to the disk again. This results in incorrect behaviour. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-23 12:34:32 +02:00
Sebastien Boeuf	f4e391922f	vmm: Remove balloon options from --memory parameter The standalone `--balloon` parameter being fully functional at this point, we can get rid of the balloon options from the --memory parameter. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-22 16:33:16 +02:00
Sebastien Boeuf	3594685279	vmm: Move balloon code from MemoryManager to DeviceManager Now that we have a new dedicated way of asking for a balloon through the CLI and the REST API, we can move all the balloon code to the device manager. This allows us to simplify the memory manager, which is already quite complex. It also simplifies the behavior of the balloon resizing command. Instead of providing the expected size for the RAM, which is complex when memory zones are involved, it now expects the balloon size. This is a much more straightforward behavior as it really resizes the balloon to the desired size. Additionally to the simplication, the benefit of this approach is that it does not need to be tied to the memory manager at all. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-22 16:33:16 +02:00
Sebastien Boeuf	1d479e5e08	vmm: Introduce new --balloon parameter This introduces a new way of defining the virtio-balloon device. Instead of going through the --memory parameter, the idea is to consider balloon as a standalone virtio device. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-22 16:33:16 +02:00
Sebastien Boeuf	28e12e9f3a	vmm, hypervisor: Fix snapshot/restore for Windows guest The snasphot/restore feature is not working because some CPU states are not properly saved, which means they can't be restored later on. First thing, we ensure the CPUID is stored so that it can be properly restored later. The code is simplified and pushed down to the hypervisor crate. Second thing, we identify for each vCPU if the Hyper-V SynIC device is emulated or not. In case it is, that means some specific MSRs will be set by the guest. These MSRs must be saved in order to properly restore the VM. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-21 19:11:03 +01:00
Rob Bradford	885ee9567b	vmm: Add support for creating virtio-watchdog The watchdog device is created through the "--watchdog" parameter. At most a single watchdog can be created per VM. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-21 16:02:39 +01:00
Michael Zhao	0b0596ef30	arch: Simplify PCI space address handling in AArch64 FDT Before Virtio-mmio was removed, we passed an optional PCI space address parameter to AArch64 code for generating FDT. The address is none if the transport is MMIO. Now Virtio-PCI is the only option, the parameter is mandatory. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-10-21 12:20:30 +01:00
Michael Zhao	2f2e10ea35	arch: Remove GICv2 Virtio-mmio is removed, now virtio-pci is the only option for virtio transport layer. We use MSI for PCI device interrupt. While GICv2, the legacy interrupt controller, doesn't support MSI. So GICv2 is not very practical for Cloud-hypervisor, we can remove it. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-10-19 14:58:48 +01:00
Sebastien Boeuf	af3c6c34c3	vmm: Remove mmio and pci differentiation Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-19 14:58:48 +01:00
Sebastien Boeuf	7cbd47a71a	vmm: Prevent KVM device fd from being unusable from VfioContainer When shutting down a VM using VFIO, the following error has been detected: vfio-ioctls/src/vfio_device.rs:312 -- Could not delete VFIO group: KvmSetDeviceAttr(Error(9)) After some investigation, it appears the KVM device file descriptor used for removing a VFIO group was already closed. This is coming from the Rust sequence of Drop, from the DeviceManager all the way down to VfioDevice. Because the DeviceManager owns passthrough_device, which is effectively a KVM device file descriptor, when the DeviceManager is dropped, the passthrough_device follows, with the effect of closing the KVM device file descriptor. Problem is, VfioDevice has not been dropped yet and it still needs a valid KVM device file descriptor. That's why the simple way to fix this issue coming from Rust dropping all resources is to make Linux accountable for it by duplicating the file descriptor. This way, even when the passthrough_device is dropped, the KVM file descriptor is closed, but a duplicated instance is still valid and owned by the VfioContainer. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-15 19:15:28 +02:00
Wei Liu	d667ed0c70	vmm: don't call notify_guest_clock_paused when Hyper-V emulation is on We turn on that emulation for Windows. Windows does not have KVM's PV clock, so calling notify_guest_clock_paused results in an error. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-10-15 19:14:25 +02:00
Sebastien Boeuf	1b9890b807	vmm: cpu: Set CPU physical bits based on user input If the user specified a maximum physical bits value through the `max_phys_bits` option from `--cpus` parameter, the guest CPUID will be patched accordingly to ensure the guest will find the right amount of physical bits. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-13 18:58:36 +02:00
Sebastien Boeuf	aec88e20d7	vmm: memory_manager: Rely on physical bits for address space size If the user provided a maximum physical bits value for the vCPUs, the memory manager will adapt the guest physical address space accordingly so that devices are not placed further than the specified value. It's important to note that if the number exceed what is available on the host, the smaller number will be picked. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-13 18:58:36 +02:00
Sebastien Boeuf	52ad78886c	vmm: Introduce new CPU option to set maximum physical bits In order to let the user choose maximum address space size, this patch introduces a new option `max_phys_bits` to the `--cpus` parameter. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-13 18:58:36 +02:00
Bo Chen	e9738a4a49	vmm: Replace the use of 'unchecked_add' with 'checked_add' The 'GuestAddress::unchecked_add' function has undefined behavior when an overflow occurs. Its alternative 'checked_add' requires use to handle the overflow explicitly. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-10-13 12:09:22 +02:00
Bo Chen	9ab2a34b40	vmm: Remove reserved 256M gaps for hotplugging memory with ACPI We are now reserving a 256M gap in the guest address space each time when hotplugging memory with ACPI, which prevents users from hotplugging memory to the maximum size they requested. We confirm that there is no need to reserve this gap. This patch removes the 'reserved gaps'. It also refactors the 'MemoryManager::start_addr' so that it is rounding-up to 128M alignment when hotplugged memory is allowed with ACPI. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-10-13 12:09:22 +02:00
Bo Chen	10f380f95b	vmm: Report no error when resizing to current memory size with ACPI We now try to create a ram region of size 0 when the requested memory size is the same as current memory size. It results in an error of `GuestMemoryRegion(Mmap(Os { code: 22, kind: InvalidInput, message: "Invalid argument" }))`. This error is not meaningful to users and we should not report it. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-10-12 08:46:38 +02:00
Bo Chen	789ee7b3e4	vmm: Support resizing memory up to and including hotplug size The start address after the hottplugged memory can be the start address of device area. Fixes: #1803 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-10-10 09:51:32 +02:00
Rob Bradford	bb5b9584d2	pci, ch-remote, vmm: Replace simple match blocks with matches! This is a new clippy check introduced in 1.47 which requires the use of the matches!() macro for simple match blocks that return a boolean. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-09 10:49:54 +02:00
Wei Liu	ed1fdd1f7d	hypervisor, arch: rename "OneRegister" and relevant code The OneRegister literally means "one (arbitrary) register". Just call it "Register" instead. There is no need to inherit KVM's naming scheme in the hypervisor agnostic code. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-10-08 08:55:10 +02:00
Sebastien Boeuf	1e3a6cb450	vmm: Simplify some of the io_uring code Small patch creating a dedicated `block_io_uring_is_supported()` function for the non-io_uring case, so that we can simplify the code in the DeviceManager. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-07 14:26:49 +02:00
Sebastien Boeuf	c02a02edfc	vmm: Allow unlink syscall for vCPU threads Without the unlink(2) syscall being allowed, Cloud-Hypervisor crashes when we remove a virtio-vsock device that has been previously added. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-06 16:05:59 +01:00
Sebastien Boeuf	6aa5e21212	vmm: device_manager: Fix PCI device unplug issues Because of the PCI refactoring that happened in the previous commit `d793cc4da3`, the ability to fully remove a PCI device was altered. The refactoring was correct, but the usage of a generic function to pass the same reference for both BusDevice, PciDevice and Any + Send + Sync causes the Arc::ptr_eq() function to behave differently than expected, as it does not match the references later in the code. That means we were not able to remove the device reference from the MMIO and/or PIO buses, which was leading to some bus range overlapping error once we were trying to add a device again to the previous range that should have been removed. Fixes #1802 Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-06 12:56:17 +02:00
Praveen Paladugu	71c435ce91	hypervisor, vmm: Introduce VmmOps trait Run loop in hypervisor needs a callback mechanism to access resources like guest memory, mmio, pio etc. VmmOps trait is introduced here, which is implemented by vmm module. While handling vcpuexits in run loop, this trait allows hypervisor module access to the above mentioned resources via callbacks. Signed-off-by: Praveen Paladugu <prapal@microsoft.com> Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-10-02 16:42:55 +01:00
Rob Bradford	2d457ab974	vmm: device_manager: Make PMEM "discard_writes" mode true CoW The PMEM support has an option called "discard_writes" which when true will prevent changes to the device from hitting the backing file. This is trying to be the equivalent of "readonly" support of the block device. Previously the memory of the device was marked as KVM_READONLY. This resulted in a trap when the guest attempted to write to it resulting a VM exit (and recently a warning). This has a very detrimental effect on the performance so instead make "discard_writes" truly CoW by mapping the memory as `PROT_READ \| PROT_WRITE` and using `MAP_PRIVATE` to establish the CoW mapping. Fixes: #1795 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-10-02 14:26:15 +02:00
Hui Zhu	c75f8b2f89	virtio-balloon: Add memory_actual_size to vm.info to show memory actual size The virtio-balloon change the memory size is asynchronous. VirtioBalloonConfig.actual of balloon device show current balloon size. This commit add memory_actual_size to vm.info to show memory actual size. Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-10-01 17:46:30 +02:00
Rob Bradford	664c3ceda6	vmm: device_manager: Warn that vhost-user self spawning is deprecated See #1724 for details. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-30 18:32:50 +02:00
Rob Bradford	0a4be7ddf5	vmm: "Cleanly" shutdown on SIGTERM Write to the exit_evt EventFD which will trigger all the devices and vCPUs to exit. This is slightly cleaner than just exiting the process as any temporary files will be removed. Fixes: #1242 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-30 18:32:16 +02:00
Bo Chen	6d30fe05e4	vmm: openapi: Add the 'iommu' and 'id' option to 'VmAddDevice' This patch adds the missing the `iommu` and `id` option for `VmAddDevice` in the openApi yaml to respect the internal data structure in the code base. Also, setting the `id` explicitly for VFIO device hotplug is required for VFIO device unplug through openAPI calls. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-09-30 08:17:44 +01:00
Julio Montes	668c563dac	vmm: openapi: fix integers format According to openAPI specification [1], the format for `integer` types can be only `int32` or `int64`, unsigned and 8-bits integers are not supported. This patch replaces `uint64` with `int64`, `uint32` with `int32` and `uint8` with `int32`. [1]: https://swagger.io/specification/#data-types Signed-off-by: Julio Montes <julio.montes@intel.com>	2020-09-29 12:55:40 -07:00
Rob Bradford	5a0d3277c8	vmm: vm: Replace \n newline character with \r This allows the CMD prompt under SAC to be used without affecting getty on Linux. Fixes: #1770 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-29 16:10:12 +02:00
Wei Liu	4ef97d8ddb	vmm: interrupts: clearly separate MsiInterruptGroup and InterruptRoute MsiInterruptGroup doesn't need to know the internal field names of InterruptRoute. Introduce two helper functions to eliminate references to irq_fd. This is done similarly to the enable and disable helper functions. Also drop the pub keyword from InterruptRoute fields. It is not needed anymore. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-09-29 13:51:35 +02:00
Praveen Paladugu	f10872e706	vmm: fix clippy warnings Signed-off-by: Praveen Paladugu <prapal@microsoft.com>	2020-09-26 14:07:12 +01:00
Praveen Paladugu	4b32252028	hypervisor, vmm: fix clippy warnings Signed-off-by: Praveen Paladugu <prapal@microsoft.com>	2020-09-26 14:07:12 +01:00
Julio Montes	c54452c08a	vmm: openapi: fix integers format According to openAPI specification[1], the format for `integer` types can be only `int32` or `int64`, unsigned integers are not supported. This patch replaces `uint64` with `int64`. [1]: https://swagger.io/specification/#data-types Signed-off-by: Julio Montes <julio.montes@intel.com>	2020-09-26 14:05:51 +01:00
Wei Liu	7e130a65ba	vmm: interrupts: adjust set_gsi_routes There is no point in manually dropping the lock for gsi_msi_routes then instantly grabbing it again in set_gsi_routes. Make set_gsi_routes take a reference to the routing hashmap instead. No functional change intended. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-09-25 17:17:35 +02:00
Sebastien Boeuf	c85e396ce5	vmm: cpu: x86: Enable MTRR feature in CPUID The MTRR feature was missing from the CPUID, which is causing the guest to ignore the MTRR settings exposed through dedicated MSRs. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-25 15:03:52 +02:00
Sebastien Boeuf	2eaf1c70c0	vmm: acpi: Advertise the correct PCI bus range Since Cloud-Hypervisor currently support one single PCI bus, we must reflect this through the MCFG table, as it advertises the first bus and the last bus available. In this case both are bus 0. This patch saves quite some time during guest kernel boot, as it prevents from checking each bus for available devices. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-23 19:03:19 +02:00
Henry Wang	961c5f2cb2	vmm: AArch64: enable VM states save/restore for AArch64 The states of GIC should be part of the VM states. This commit enables the AArch64 VM states save/restore by adding save/restore of GIC states. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	3ea4a0797d	vmm: seccomp: unify AArch64 and x86_64 FTRUNCATE syscall The definition of libc::SYS_ftruncate on AArch64 is different from that on x86_64. This commit unifies the previously hard-coded syscall number for AArch64. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	48544e4e82	vmm: seccomp: whitelist `KVM_GET_REG_LIST` in seccomp `KVM_GET_REG_LIST` ioctl is needed in save/restore AArch64 vCPU. Therefore we whitelist this ioctl in seccomp. Also this commit unifies the `SYS_FTRUNCATE` syscall for x86_64 and AArch64. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	c6b47d39e0	vmm: refactor vCPU save/restore code in restoring VM Similarly as the VM booting process, on AArch64 systems, the vCPUs should be created before the creation of GIC. This commit refactors the vCPU save/restore code to achieve the above-mentioned restoring order. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	970a5a410d	vmm: decouple vCPU init from `configure_vcpus` Since calling `KVM_GET_ONE_REG` before `KVM_VCPU_INIT` will result in an error: Exec format error (os error 8). This commit decouples the vCPU init process from `configure_vcpus`. Therefore in the process of restoring the vCPUs, these vCPUs can be initialized separately before started. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	47e65cd341	vmm: AArch64: add methods to get saved vCPU states The construction of `GICR_TYPER` register will need vCPU states. Therefore this commit adds methods to extract saved vCPU states from the cpu manager. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	381d0b4372	devices: remove the migration traits for the `Gic` struct Unlike x86_64, the "interrupt_controller" in the device manager for AArch64 is only a `Gic` object that implements the `InterruptController` to provide the interrupt delivery service. This is not the real GIC device so that we do not need to save its states. Also, we do not need to insert it to the device_tree. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	7ddcad1d8b	arch: AArch64: add a field `gicr_typers` for GIC implementations The value of GIC register `GICR_TYPER` is needed in restoring the GIC states. This commit adds a field in the GIC device struct and a method to construct its value. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	dcf6d9d731	device_manager: AArch64: add a field to set/get GIC device entity In AArch64 systems, the state of GIC device can only be retrieved from `KVM_GET_DEVICE_ATTR` ioctl. Therefore to implement saving/restoring the GIC states, we need to make sure that the GIC object (either the file descriptor or the device itself) can be extracted after the VM is started. This commit refactors the code of GIC creation by adding a new field `gic_device_entity` in device manager and methods to set/get this field. The GIC object can be therefore saved in the device manager after calling `arch::configure_system`. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	e7acbcc184	arch: AArch64: support saving RDIST pending tables into guest RAM This commit adds a function which allows to save RDIST pending tables to the guest RAM, as well as unit test case for it. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	29ce3076c2	tests: AArch64: Add unit test cases for accessing GIC registers This commit adds the unit test cases for getting/setting the GIC distributor, redistributor and ICC registers. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	9dd188a8e8	tests: AArch64: Add unit test cases for vCPU save/restore Adds 3 more unit test cases for AArch64: save_restore_core_regs save_restore_system_regs *get_set_mpstate Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Henry Wang	e3d45be6f7	AArch64: Preparation for vCPU save/restore This commit ports code from firecracker and refactors the existing AArch64 code as the preparation for implementing save/restore AArch64 vCPU, including: 1. Modification of `arm64_core_reg` macro to retrive the index of arm64 core register and implemention of a helper to determine if a register is a system register. 2. Move some macros and helpers in `arch` crate to the `hypervisor` crate. 3. Added related unit tests for above functions and macros. Signed-off-by: Henry Wang <Henry.Wang@arm.com>	2020-09-23 12:37:25 +01:00
Josh Soref	5c3f4dbe6f	ch: Fix various misspelled words Misspellings were identified by https://github.com/marketplace/actions/check-spelling * Initial corrections suggested by Google Sheets * Additional corrections by Google Chrome auto-suggest * Some manual corrections Signed-off-by: Josh Soref <jsoref@users.noreply.github.com>	2020-09-23 08:59:31 +01:00
Jiangbo Wu	22a2a99e5f	acpi: Add hotplug numa node virtio-mem device would use 'VIRTIO_MEM_F_ACPI_PXM' to add memory to NUMA node, which MUST be existed, otherwise it will be assigned to node id 0, even if user specify different node id. According ACPI spec about Memory Affinity Structure, system hardware supports hot-add memory region using 'Hot Pluggable \| Enabled' flags. Signed-off-by: Jiangbo Wu <jiangbo.wu@intel.com>	2020-09-22 13:11:39 +02:00
Jiangbo Wu	223189c063	mm: Apply zone's property instread of global config Apply memory zone's property for associated virtio-mem regions. Signed-off-by: Jiangbo Wu <jiangbo.wu@intel.com>	2020-09-22 09:56:37 +02:00
Jiangbo Wu	80be8ac0dc	mm: Apply memory policy for virtio-mem region Use zone.host_numa_node to create memory zone, so that memory zone can apply memory policy in according with host numa node ID Signed-off-by: Jiangbo Wu <jiangbo.wu@intel.com>	2020-09-22 09:56:37 +02:00
Sebastien Boeuf	7c346c3844	vmm: Kill vhost-user self-spawned process on failure If after the creation of the self-spawned backend, the VMM cannot create the corresponding vhost-user frontend, the VMM must kill the freshly spawned process in order to ensure the error propagation can happen. In case the child process would still be around, the VMM cannot return the error as it waits onto the child to terminate. This should help us identify when self-spawned failures are caused by a connection being refused between the VMM and the backend. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-18 17:26:25 +01:00
Sebastien Boeuf	555c5c5d9c	vmm: Add missing syscalls to signal thread When the VMM is terminated by receiving a SIGTERM signal, the signal handler thread must be able to invoke ioctl(TCGETS) and ioctl(TCSETS) without error. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-18 13:40:10 +01:00
Rob Bradford	41a9b1adef	vmm: Add missing syscall to vCPU thread Fixes: #1717 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-18 13:40:10 +01:00
Sebastien Boeuf	1e1a50ef70	vmm: Update memory configuration upon virtio-mem resizing Based on all the preparatory work achieved through previous commits, this patch updates the 'hotplugged_size' field for both MemoryConfig and MemoryZoneConfig structures when either the whole memory is resized, or simply when a memory zone is resized. This fixes the lack of support for rebooting a VM with the right amount of memory plugged in. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	de2b917f55	vmm: Add hotplugged_size to VirtioMemZone Adding a new field to VirtioMemZone structure, as it lets us associate with a particular virtio-mem region the amount of memory that should be plugged in at boot. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	3faf8605f3	vmm: Group virtio-mem fields under a dedicated structure This patch simplifies the code as we have one single Option for the VirtioMemZone. This also prepares for storing additional information related to the virtio-mem region. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	4e1b78e1ff	vmm: Add 'hotplugged_size' to memory parameters Add the new option 'hotplugged_size' to both --memory-zone and --memory parameters so that we can let the user specify a certain amount of memory being plugged at boot. This is also part of making sure we can store the virtio-mem size over a reboot of the VM. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Hui Zhu	33a1e37c35	virtio-devices: mem: Allow for an initial size This commit gives the possibility to create a virtio-mem device with some memory already plugged into it. This is preliminary work to be able to reboot a VM with the virtio-mem region being already resized. Signed-off-by: Hui Zhu <teawater@antfin.com> Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	8b5202aa5a	vmm: Always add virtio-mem region upon VM creation Now that e820 tables are created from the 'boot_guest_memory', we can simplify the memory manager code by adding the virtio-mem regions when they are created. There's no need to wait for the first hotplug to insert these regions. This also anticipates the need for starting a VM with some memory already plugged into the virtio-mem region. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	66fc557015	vmm: Store boot guest memory and use it for boot sequence In order to differentiate the 'boot' memory regions from the virtio-mem regions, we store what we call 'boot_guest_memory'. This is useful to provide the adequate list of regions to the configure_system() function as it expects only the list of regions that should be exposed through the e820 table. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	1798ed8194	vmm: virtio-mem: Enforce alignment and size requirements The virtio-mem driver is generating some warnings regarding both size and alignment of the virtio-mem region if not based on 128MiB: The alignment of the physical start address can make some memory unusable. The alignment of the physical end address can make some memory unusable. For these reasons, the current patch enforces virtio-mem regions to be 128MiB aligned and checks the size provided by the user is a multiple of 128MiB. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	eb7b923e22	vmm: Create virtio-mem device with appropriate NUMA node Now that virtio-mem device accept a guest NUMA node as parameter, we retrieve this information from the list of NUMA nodes. Based on the memory zone associated with the virtio-mem device, we obtain the NUMA node identifier, which we provide to the virtio-mem device. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	dcedd4cded	virtio-devices: virtio-mem: Add NUMA support Implement support for associating a virtio-mem device with a specific guest NUMA node, based on the ACPI proximity domain identifier. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	0658559880	vmm: memory_manager: Rename 'use_zones' with 'user_provided_zones' This brings more clarity on the meaning of this boolean. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	775f3346e3	vmm: Rename 'virtiomem' to 'virtio_mem' For more consistency and help reading the code better, this commit renames all 'virtiomem' variables into 'virtio_mem'. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	015c78411e	vmm: Add a 'resize-zone' action to the API actions Implement a new VM action called 'resize-zone' allowing the user to resize one specific memory zone at a time. This relies on all the preliminary work from the previous commits to resize each virtio-mem device independently from each others. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	141df701dd	vmm: memory_manager: Make virtiomem_resize function generic By adding a new parameter 'id' to the virtiomem_resize() function, we prepare this function to be usable for both global memory resizing and memory zone resizing. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	34331d3e72	vmm: memory_manager: Fix virtio-mem resize It's important to return the region covered by virtio-mem the first time it is inserted as the device manager must update all devices with this information. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	adc59a6f15	vmm: memory_manager: Create one virtio-mem per memory zone Based on the previous code changes, we can now update the MemoryManager code to create one virtio-mem region and resizing handler per memory zone. This will naturally create one virtio-mem device per memory zone from the DeviceManager's code which has been previously updated as well. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	c645a72c17	vmm: Add 'hotplug_size' to memory zones In anticipation for resizing support of an individual memory zone, this commit introduces a new option 'hotplug_size' to '--memory-zone' parameter. This defines the amount of memory that can be added through each specific memory zone. Because memory zone resize is tied to virtio-mem, make sure the user selects 'virtio-mem' hotplug method, otherwise return an error. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	30ff7e108f	vmm: Prepare code to accept multiple virtio-mem devices Both MemoryManager and DeviceManager are updated through this commit to handle the creation of multiple virtio-mem devices if needed. For now, only the framework is in place, but the behavior remains the same, which means only the memory zone created from '--memory' generates a virtio-mem region that can be used for resize. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Sebastien Boeuf	b173b6c5b4	vmm: Create a MemoryZone structure In order to anticipate the need for storing memory regions along with virtio-mem information for each memory zone, we create a new structure MemoryZone that will replace Vec<Arc<GuestRegionMmap>> in the hash map MemoryZones. This makes thing more logical as MemoryZones becomes a list of MemoryZone sorted by their identifier. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 19:20:04 +02:00
Rob Bradford	27c28fa3b0	vmm, arch: Enable KVM HyperV support Inject CPUID leaves for advertising KVM HyperV support when the "kvm_hyperv" toggle is enabled. Currently we only enable a selection of features required to boot. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-16 16:08:01 +01:00
Rob Bradford	da642fcf7f	hypervisor: Add "HyperV" exit to list of KVM exits Currently we don't need to do anything to service these exits but when the synthetic interrupt controller is active an exit will be triggered to notify the VMM of details of the synthetic interrupt page. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-16 16:08:01 +01:00
Rob Bradford	5495ab7415	vmm: Add "kvm_hyperv" toggle to "--cpus" This turns on the KVM HyperV emulation. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-16 16:08:01 +01:00
Sebastien Boeuf	b3435d51d9	vmm: cpu: Add missing io_uring syscalls to vCPU threads Some of the io_uring setup happens upon activation of the virtio-blk device, which is initially triggered through an MMIO VM exit. That's why the vCPU threads must authorize io_uring related syscalls. This commit ensures the virtio-blk io_uring implementation can be used along with the seccomp filters enabled. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-16 11:59:47 +02:00
Bo Chen	9682d74763	vmm: seccomp: Add seccomp filters for signal_handler worker thread This patch covers the last worker thread with dedicated secomp filters. Fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-09-11 07:42:31 +02:00
Bo Chen	2612a6df29	vmm: seccomp: Add seccomp filters for the vcpu worker thread Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-09-11 07:42:31 +02:00
Rob Bradford	d793cc4da3	vmm: device_manager: Extract common PCI code Extract common code for adding devices to the PCI bus into its own function from the VFIO and VIRTIO code paths. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-11 07:33:18 +02:00
Rob Bradford	15025d71b1	devices, vm-device: Move BusDevice and Bus into vm-device This removes the dependency of the pci crate on the devices crate which now only contains the device implementations themselves. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-09-10 09:35:38 +01:00
Bo Chen	3c923f0727	virtio-devices: seccomp: Add seccomp filters for virtio_vsock thread This patch enables the seccomp filters for the virtio_vsock worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-09-09 17:04:39 +01:00
Bo Chen	1175fa2bc7	virtio-devices: seccomp: Add seccomp filters for blk_io_uring thread This patch enables the seccomp filters for the block_io_uring worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-09-09 17:04:39 +01:00
Sebastien Boeuf	e15dba2925	vmm: Rename NUMA option 'id' into 'guest_numa_id' The goal of this commit is to rename the existing NUMA option 'id' with 'guest_numa_id'. This is done without any modification to the way this option behaves. The reason for the rename is caused by the observation that all other parameters with an option called 'id' expect a string to be provided. Because in this particular case we expect a u32 representing a proximity domain from the ACPI specification, it's better to name it with a more explicit name. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Sebastien Boeuf	1970ee89da	main, vmm: Remove guest_numa_node option from memory zones The way to describe guest NUMA nodes has been updated through previous commits, letting the user describe the full NUMA topology through the --numa parameter (or NumaConfig). That's why we can remove the deprecated and unused 'guest_numa_node' option. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Sebastien Boeuf	f21c04166a	vmm: Move NUMA node list creation to Vm structure Based on the previous changes introducing new options for both memory zones and NUMA configuration, this patch changes the behavior of the NUMA node definition. Instead of relying on the memory zones to define the guest NUMA nodes, everything goes through the --numa parameter. This allows for defining NUMA nodes without associating any particular memory range to it. And in case one wants to associate one or multiple memory ranges to it, the expectation is to describe a list of memory zone through the --numa parameter. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Sebastien Boeuf	dc42324351	vmm: Add 'memory_zones' option to NumaConfig This new option provides a new way to describe the memory associated with a NUMA node. This is the first step before we can remove the 'guest_numa_node' option from the --memory-zone parameter. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Sebastien Boeuf	5d7215915f	vmm: memory_manager: Store a list of memory zones Now that we have an identifier per memory zone, and in order to keep track of the memory regions associated with the memory zones, we create and store a map referencing list of memory regions per memory zone ID. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Sebastien Boeuf	3ff82b4b65	main, vmm: Add mandatory id to memory zones In anticipation for allowing memory zones to be removed, but also in anticipation for refactoring NUMA parameter, we introduce a mandatory 'id' option to the --memory-zone parameter. This forces the user to provide a unique identifier for each memory zone so that we can refer to these. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-07 07:37:14 +02:00
Samuel Ortiz	e5ce6dc43c	vmm: cpu: Warn if the guest is trying to access unregistered IO ranges Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-09-04 14:39:58 +02:00
Sebastien Boeuf	c0d0d23932	vmm: acpi: Introduce SLIT for NUMA nodes distances By introducing the SLIT (System Locality Distance Information Table), we provide the guest with the distance between each node. This lets the user describe the NUMA topology with a lot of details so that slower memory backing the VM can be exposed as being further away from other nodes. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 18:09:01 +02:00
Sebastien Boeuf	9548e7e857	vmm: Update NUMA node distances internally Based on the NumaConfig which now provides distance information, we can internally update the list of NUMA nodes with the exact distances they should be located from other nodes. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 18:09:01 +02:00
Sebastien Boeuf	a5a29134ca	vmm: Extend --numa parameter with NUMA node distances By introducing 'distances' option, we let the user describe a list of destination NUMA nodes with their associated distances compared to the current node (defined through 'id'). Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 18:09:01 +02:00
Sebastien Boeuf	629befdb4a	vmm: acpi: Add CPUs to NUMA nodes Based on the list of CPUs related to each NUMA node, Processor Local x2APIC Affinity structures are created and included into the SRAT table. This describes which CPUs are part of each node. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 15:25:00 +02:00
Sebastien Boeuf	db28db8567	vmm: Update NUMA nodes based on NumaConfig Relying on the list of CPUs defined through the NumaConfig, this patch will update the internal list of CPUs attached to each NUMA node. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 15:25:00 +02:00
Sebastien Boeuf	42f963d6f2	main, vmm: Add new --numa parameter Through this new parameter, we give users the opportunity to specify a set of CPUs attached to a NUMA node that has been previously created from the --memory-zone parameter. This parameter will be extended in the future to describe the distance between multiple nodes. For instance, if a user wants to attach CPUs 0, 1, 2 and 6 to a NUMA node, here are two different ways of doing so: Either ./cloud-hypervisor ... --numa id=0,cpus=0-2:6 Or ./cloud-hypervisor ... --numa id=0,cpus=0:1:2:6 Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 15:25:00 +02:00
Sebastien Boeuf	65a23c6fc6	vmm: acpi: Create the SRAT table The SRAT table (System Resource Affinity Table) is needed to describe NUMA nodes and how memory ranges and CPUs are attached to them. For now it simply attaches a list of Memory Affinity structures based on the list of NUMA nodes created from the VMM. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 14:11:49 +02:00
Sebastien Boeuf	cf81254a8d	vmm: memory_manager: Create a NUMA node list Based on the 'guest_numa_node' option, we create and store a list of NUMA nodes in the MemoryManager. The point being to associate a list of memory regions to each node, so that we can later create the ACPI tables with the proper memory range information. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 14:11:49 +02:00
Sebastien Boeuf	768dbd1fb0	vmm: Add 'guest_numa_node' option to 'memory-zone' With the introduction of this new option, the user will be able to describe if a particular memory zone should belong to a specific NUMA node from a guest perspective. For instance, using '--memory-zone size=1G,guest_numa_node=2' would let the user describe that a memory zone of 1G in the guest should be exposed as being associated with the NUMA node 2. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 14:11:49 +02:00
Sebastien Boeuf	274c001eab	vmm: Use u32 instead of u64 for host_numa_node option Given that ACPI uses u32 as the type for the Proximity Domain, we can use u32 instead of u64 as the type for 'host_numa_node' option. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-09-01 13:29:42 +02:00
Michael Zhao	a95b6bbd8b	vmm: Add seccomp rules for starting vhost-user-net backend on AArch64 Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-08-31 08:19:23 +02:00
Hui Zhu	f7b3581645	cloud-hypervisor.yaml: MemoryConfig: Add balloon_size "struct MemoryConfig" has balloon_size but not in MemoryConfig of cloud-hypervisor.yaml. This commit adds it. Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-08-28 09:58:39 +02:00
Sebastien Boeuf	a8a9e61c3d	vmm: memory_manager: Allow host NUMA for RAM backed files Let's narrow down the limitation related to mbind() by allowing shared mappings backed by a file backed by RAM. This leaves the restriction on only for mappings backed by a regular file. With this patch, host NUMA node can be specified even if using vhost-user devices. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-27 08:39:38 -07:00
Sebastien Boeuf	1b4591aecc	vmm: memory_manager: Apply NUMA policy to memory zones Relying on the new option 'host_numa_node' from the 'memory-zone' parameter, the user can now define which NUMA node from the host should be used to back the current memory zone. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-27 08:39:38 -07:00
Sebastien Boeuf	e6f585a31c	vmm: Add 'host_numa_nodes' option to memory zones Since memory zones have been introduced, it is now possible for a user to specify multiple backends for the guest RAM. By adding a new option 'host_numa_node' to the 'memory-zone' parameter, we allow the guest RAM to be backed by memory that might come from a specific NUMA node on the host. The option expects a node identifier, specifying which NUMA node should be used to allocate the memory associated with a specific memory zone. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-27 08:39:38 -07:00
Sebastien Boeuf	ad5d0e4713	vmm: Remove 'mergeable' from memory zones The flag 'mergeable' should only apply to the entire guest RAM, which is why it is removed from the MemoryZoneConfig as it is defined as a global parameter at the MemoryConfig level. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-27 07:26:49 +02:00
Sebastien Boeuf	89e7774b96	vmm: openapi: Don't expect cmdline to always be there The 'cmdline' parameter should not be required as it is not needed when the 'kernel' parameter is the rust-hypervisor-fw, which means the kernel and the associated command line will be found from the EFI partition. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:49:05 +02:00
Sebastien Boeuf	e8149380b7	vmm: memory_manager: Factorize memory regions creation Factorize the codepath between simple memory and multiple memory zones. This simplifies the way regions are memory mapped, as everything relies on the same codepath. This is performed by creating a memory zone on the fly for the specific use case where --memory is used with size being different from 0. Internally, the code can rely on memory zones to create the memory regions forming the guest memory. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	c58dd761f4	vmm: Remove 'file' option from MemoryConfig After the introduction of user defined memory zones, we can now remove the deprecated 'file' option from --memory parameter. This makes this parameter simpler, letting more advanced users define their own custom memory zones through the dedicated parameter. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	5bf7113768	vmm: memory_manager: Remove restrictions about snapshot/restore User defined memory regions can now support being snapshot and restored, therefore this commit removes the restrictions that were applied through earlier commit. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	2583d572fc	vmm: memory_manager: Simplify how to restore memory regions By factorizing a lot of code into create_ram_region(), this commit achieves the simplification of the restore codepath. Additionally, it makes user defined memory zones compatible with snapshot/restore. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	b14c861c6f	vmm: memory_manager: Store memory regions content only when necessary First thing, this patch introduces a new function to identify if a file descriptor is linked to any hard link on the system. This can let the VMM know if the file can be accessed by the user, or if the file will be destroyed as soon as the VMM releases the file descriptor. Based on this information, and associated with the knowledge about the region being MAP_SHARED or not, the VMM can now decide to skip the copy of the memory region content. If the user has access to the file from the filesystem, and if the file has been mapped as MAP_SHARED, we can consider the guest memory region content to be present in this file at any point in time. That's why in this specific case, there's no need for performing the copy of the memory region content into a dedicated file. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	d1ce52f3a8	vmm: memory_manager: Make backing file from snapshot optional Let's not assume that a backing file is going to be the result from a snapshot for each memory region. These regions might be backed by a file on the host filesystem (not a temporary file in host RAM), which means they don't need to be copied and stored into dedicated files. That's why this commit prepares for further changes by introducing an optional PathBuf associated with the snapshot of each memory region. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	871138d5cc	vm-migration: Make snapshot() mutable There will be some cases where the implementation of the snapshot() function from the Snapshottable trait will require to modify some internal data, therefore we make this possible by updating the trait definition with snapshot(&mut self). Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	c13721fdbd	vmm: memory_manager: Handle user defined memory zones In case the memory size is 0, this means the user defined memory zones are used as a way to specify how to back the guest memory. This is the first step in supporting complex use cases where the user can define exactly which type of memory from the host should back the memory from the guest. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	7cd3867e2c	vmm: memory_manager: Provide file offset through create_ram_region() In anticipation for the need to map part of a file with the function create_ram_region(), it is extended to accept a file offset as argument. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	59d4a56ab7	vmm: memory_manager: Don't truncate backing file In case the provided backing file is an actual file and not a directory, we should not truncate it, as we expect the file to already be the right size. This change will be important once we try to map the same file through multiple memory mappings. We can't let the file be truncated as the second mapping wouldn't work properly. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	be475ddc22	main, vmm: Let the user define distincts memory zones Introducing a new CLI option --memory-zone letting the user specify custom memory zones. When this option is present, the --memory size must be explicitly set to 0. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Sebastien Boeuf	d25ec66bb6	vmm: memory_manager: Simplify start_addr() Small simplification for the function calculating the start address. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-25 16:43:10 +02:00
Anatol Belski	12212d2966	pci: device_manager: Remove hardcoded I/O port assignment It is otherwise seems to be able to cause resource conflicts with Windows APCI_HAL. The OS might do a better job on assigning resources to this device, withouth them to be requested explicitly. 0xcf8 and 0xcfc are only what is certainly needed for the PCI device enumeration. Signed-off-by: Anatol Belski <anatol.belski@microsoft.com>	2020-08-25 09:00:06 +02:00
Michael Zhao	afc98a5ec9	vmm: Fix AArch64 clippy warnings of vmm and other crates Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-08-24 10:59:08 +02:00
Muminul Islam	92b4499c1e	vmm, hypervisor: Add vmstate to snapshot and restore path Signed-off-by: Muminul Islam <muislam@microsoft.com>	2020-08-24 08:48:15 +02:00
Bo Chen	02d87833f0	virtio-devices: seccomp: Add seccomp filters for vhost_blk thread This patch enables the seccomp filters for the vhost_blk worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-19 08:33:58 +02:00
Bo Chen	896b9a1d4b	virtio-devices: seccomp: Add seccomp filter for vhost_net_ctl thread This patch enables the seccomp filters for the vhost_net_ctl worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-19 08:33:58 +02:00
Bo Chen	02d63149fe	virtio-devices: seccomp: Add seccomp filters for vhost_fs thread This patch enables the seccomp filters for the vhost_fs worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-19 08:33:58 +02:00
Bo Chen	c82ded8afa	virtio-devices: seccomp: Add seccomp filters for balloon thread This patch enables the seccomp filters for the balloon worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-19 08:33:58 +02:00
Bo Chen	c460178723	virtio-devices: seccomp: Add seccomp filters for mem thread This patch enables the seccomp filters for the mem worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-19 08:33:58 +02:00
Bo Chen	4539236690	virtio-devices: seccomp: Add seccomp filters for iommu thread This patch enables the seccomp filters for the iommu worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-17 21:08:49 +02:00
Anatol Belski	eba42c392f	devices: acpi: Add UID to devices with common HID Some OS might check for duplicates and bail out, if it can't create a distinct mapping. According to ACPI 5.0 section 6.1.12, while _UID is optional, it becomes required when there are multiple devices with the same _HID. Signed-off-by: Anatol Belski <ab@php.net>	2020-08-14 08:52:02 +02:00
Sebastien Boeuf	bdef54ead6	vmm: Add brk syscall to the API thread The brk syscall is not always called as the system might not need it. But when it's needed from the API thread, this causes the thread to terminate as it is not part of the authorized list of syscalls. This should fix some sporadic failures on the CI with the musl build. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-11 15:04:21 +01:00
Jose Carlos Venegas Munoz	90acb01bad	vmm: seccomp: add mprotect to API thread filter Add mprotect to API thread rules. Prevent the VMM is killed when it is used. Signed-off-by: Jose Carlos Venegas Munoz <jose.carlos.venegas.munoz@intel.com>	2020-08-05 21:35:21 +01:00
Bo Chen	dc71d2765a	virtio-devices: seccomp: Add seccomp filters for pmem thread This patch enables the seccomp filters for the pmem worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-05 08:13:31 +01:00
Bo Chen	d77977536d	virtio-devices: seccomp: Add seccomp filters for net thread This patch enables the seccomp filters for the net worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-05 08:13:31 +01:00
Bo Chen	276df6b71c	virtio-devices: seccomp: Add seccomp filters for console thread This patch enables the seccomp filters for the console worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-05 08:13:31 +01:00
Bo Chen	a426221167	virtio-devices: seccomp: Add seccomp filters for rng thread This patch enables the seccomp filters for the rng worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-05 08:13:31 +01:00
Bo Chen	704edd544c	virtio-devices: seccomp: Add seccomp_filter module This patch added the seccomp_filter module to the virtio-devices crate by taking reference code from the vmm crate. This patch also adds allowed-list for the virtio-block worker thread. Partially fixes: #925 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-04 11:40:49 +02:00
Bo Chen	ff7ed8f628	vmm: Propagate the SeccompAction value to the Vm struct constructor This patch propagates the SeccompAction value from main to the Vm struct constructor (i.e. Vm::new_from_memory_manager), so that we can use it to construct the DeviceManager and CpuManager struct for controlling the behavior of the seccomp filters for vcpu/virtio-device worker threads. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-04 11:40:49 +02:00
Bo Chen	8e74637ebb	main, vmm: seccomp: Add the '--seccomp log' option This patch extends the CLI option '--seccomp' to accept the 'log' parameter in addition 'true/false'. It also refactors the vmm::seccomp_filters module to support both "SeccompAction::Trap" and "SeccompAction::Log". Fixes: #1180 Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-04 11:40:49 +02:00
Bo Chen	b41884a406	main, vmm: seccomp: Use SeccompAction instead of SeccompLevel This patch replaces the usage of 'SeccompLevel' with 'SeccompAction', which is the first step to support the 'log' action over system calls that are not on the allowed list of seccomp filters. Signed-off-by: Bo Chen <chen.bo@intel.com>	2020-08-04 11:40:49 +02:00
Sebastien Boeuf	8f0bf82648	io_uring: Add new feature gate By adding a new io_uring feature gate, we let the user the possibility to choose if he wants to enable the io_uring improvements or not. Since the io_uring feature depends on the availability on recent host kernels, it's better if we leave it off for now. As soon as our CI will have support for a kernel 5.6 with all the features needed from io_uring, we'll enable this feature gate permanently. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-03 14:15:01 +01:00
Sebastien Boeuf	917027c55b	vmm: Rely on virtio-blk io_uring when possible In case the host supports io_uring and the specific io_uring options needed, the VMM will choose the asynchronous version of virtio-blk. This will enable better I/O performances compared to the default synchronous version. This is also important to note the VMM won't be able to use the asynchronous version if the backend image is in QCOW format. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-08-03 14:15:01 +01:00
Praveen Paladugu	afa8ecc90c	vmm: add validation for network parameters Signed-off-by: Praveen Paladugu <prapal@microsoft.com>	2020-07-31 09:07:12 +02:00
Wei Liu	a52b614a61	vmm: device_manager: console input should be only consumed by one device Cloud Hypervisor allows either the serial or virtio console to output to TTY, but TTY input is pushed to both. This is not correct. When Linux guest is configured to spawn TTYs on both ttyS0 and hvc0, the user effectively issues the same commands twice in different TTYs. Fix this by only direct input to the one choice that is using host side TTY. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-30 18:05:01 +02:00
Wei Liu	5ed794a44c	vmm: device_manager: rename console_input to virtio_console_input Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-30 18:05:01 +02:00
Wei Liu	3e68867bb7	vmm: device_manager: eliminate KvmMsiInterruptManager from the new function The logic to create an MSI interrupt manager is applicable to Hyper-V as well. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-30 08:00:33 +02:00
Wei Liu	218ec563fc	vmm: fix warnings when KVM is not enabled Some imports are only used by KVM. Some variables and code become dead or unused when KVM is not enabled. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-28 21:08:39 +01:00
Jianyong Wu	d24b110519	seccomp: AArch64: Add SYS_unlinkat to seccomp whitelist This commit fixes an "Bad syscall" error when shutting down the VM on AArch64 by adding the SYS_unlinkat syscall to the seccomp whitelist. Signed-off-by: Jianyong Wu <jianyong.wu@arm.com>	2020-07-27 07:25:07 +00:00
Rob Bradford	9ae44aeada	vmm: acpi_tables: Fix PM timer I/O port width Ensure that the width of the I/O port is correctly set to 32-bits in the generic address used for the X_PM_TMR_BLK. Do this by type parameterising GenericAddress::io_port_address() fuction. TEST=Boot with clocksource=acpi_pm and observe no errors in the dmesg. Fixes: #1496 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-07-23 17:48:22 +02:00
Rob Bradford	aae5d988e1	devices: vmm: Add ACPI PM timer This is a counter exposed via an I/O port that runs at 3.579545MHz. Here we use a hardcoded I/O and expose the details through the FADT table. TEST=Boot Linux kernel and see the following in dmesg: [ 0.506198] clocksource: acpi_pm: mask: 0xffffff max_cycles: 0xffffff, max_idle_ns: 2085701024 ns Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-07-23 13:10:21 +01:00
Wei Liu	f03afea0d6	device_manager: document unsafe block in add_vfio_device It is not immediately obvious why the conversion is safe. Document the safety guarantee. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-21 17:13:10 +01:00
Samuel Ortiz	be51ea250d	device_manager: Simplify the passthrough internal API We store the device passthrough handler, so we should use it through our internal API and only carry the passed through device configuration. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-21 17:20:25 +02:00
Michael Zhao	ddf1b76906	hypervisor: Refactor create_passthrough_device() for generic type Changed the return type of create_passthrough_device() to generic type hypervisor::Device. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-21 16:22:02 +02:00
Michael Zhao	e3e771727a	arch: Refactor GIC code to seperate KVM specific code Shrink GICDevice trait to contain hypervisor agnostic API's only, which are used in generating FDT. Move all KVM specific logic into KvmGICDevice trait. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-21 16:22:02 +02:00
Michael Zhao	3e051e7b2c	arch, vmm: Enable initramfs on AArch64 Ported Firecracker commit 144b6c. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-20 14:20:53 +01:00
Wei Liu	e1af251c9f	vmm, hypervisor: adjust set_gsi_routing / set_gsi_routes Make set_gsi_routing take a list of IrqRoutingEntry. The construction of hypervisor specific structure is left to set_gsi_routing. Now set_gsi_routes, which is part of the interrupt module, is only responsible for constructing a list of routing entries. This further splits hypervisor specific code from hypervisor agnostic code. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-20 07:32:32 +02:00
Wei Liu	d484a3383c	vmm: device_manager: introduce add_passthrough_device It calls add_vfio_device on KVM or returns an error when not running on KVM. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-17 20:21:39 +02:00
Wei Liu	821892419c	vmm: device_manager: use generic names for passthrough device Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-17 20:21:39 +02:00
Wei Liu	ff8d7bfe83	hypervisor: add create_passthrough_device call to Vm trait That function is going to return a handle for passthrough related operations. Move create_kvm_device code there. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-17 20:21:39 +02:00
Wei Liu	c08d2b2c70	device_manager: avoid manipulating MemoryRegion fields directly Hyper-V may have different field names. Use make_user_memory_region instead. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-16 15:56:03 +02:00
Wei Liu	d80e383dbb	arch: move test cases to vmm crate This saves us from adding a "kvm" feature to arch crate merely for the purpose of running tests. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-15 17:21:07 +02:00
Wei Liu	598eaf9f86	vmm: use hypervisor::new in test_vm Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-15 17:21:07 +02:00
Sebastien Boeuf	a5c4f0fc6f	arch, vmm: Add e820 entry related to SGX EPC region SGX expects the EPC region to be reported as "reserved" from the e820 table. This patch adds a new entry to the table if SGX is enabled. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Sebastien Boeuf	e10d9b13d4	arch, hypervisor, vmm: Patch CPUID subleaves to expose EPC sections The support for SGX is exposed to the guest through CPUID 0x12. KVM passes static subleaves 0 and 1 from the host to the guest, without needing any modification from the VMM itself. But SGX also relies on dynamic subleaves 2 through N, used for describing each EPC section. This is not handled by KVM, which means the VMM is in charge of setting each subleaf starting from index 2 up to index N, depending on the number of EPC sections. These subleaves 2 through N are not listed as part of the supported CPUID entries from KVM. But it's important to set them as long as index 0 and 1 are present and indicate that SGX is supported. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Sebastien Boeuf	1603786374	vmm: Pass MemoryManager through CpuManager creation Instead of passing the GuestMemoryMmap directly to the CpuManager upon its creation, it's better to pass a reference to the MemoryManager. This way we will be able to know if SGX EPC region along with one or multiple sections are present. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Sebastien Boeuf	2b06ce0ed4	vmm: Add EPC device to ACPI tables The SGX EPC region must be exposed through the ACPI tables so that the guest can detect its presence. The guest only get the full range from ACPI, as the specific EPC sections are directly described through the CPUID of each vCPU. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Sebastien Boeuf	84cf12d86a	arch, vmm: Create SGX virtual EPC sections from MemoryManager Based on the presence of one or multiple SGX EPC sections from the VM configuration, the MemoryManager will allocate a contiguous block of guest address space to hold the entire EPC region. Within this EPC region, each EPC section is memory mapped. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Sebastien Boeuf	d9244e9f4c	vmm: Add option for enabling SGX EPC regions Introducing the new CLI option --sgx-epc along with the OpenAPI structure SgxEpcConfig, so that a user can now enable one or multiple SGX Enclave Page Cache sections within a contiguous region from the guest address space. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-07-15 15:08:56 +02:00
Michael Zhao	cce6237536	pci: Enable GSI routing (MSI type) for AArch64 In this commit we saved the BDF of a PCI device and set it to "devid" in GSI routing entry, because this field is mandatory for GICv3-ITS. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-14 14:34:54 +01:00
Michael Zhao	f2e484750a	arch: aarch64: Add PCIe node in FDT for AArch64 Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-14 14:34:54 +01:00
Michael Zhao	17057a0dd9	vmm: Fix build errors with "pci" feature on AArch64 Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-14 14:34:54 +01:00
Rob Bradford	4963e37dc8	qcow, virtio-devices: Break cyclic dependency Move the definition of RawFile from virtio-devices crate into qcow crate. All the code that consumes RawFile also already depends on the qcow crate for image file type detection so this change breaks the need for the qcow crate to depend on the very large virtio-devices crate. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-07-10 17:47:31 +02:00
Hui Zhu	800220acbb	virtio-balloon: Store the balloon size to support reboot This commit store balloon size to MemoryConfig. After reboot, virtio-balloon can use this size to inflate back to the size before reboot. Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-07-07 17:25:13 +01:00
Hui Zhu	8ffbc3d031	vmm: api: ch-remote: Add balloon to VmResizeData Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-07-07 17:25:13 +01:00
Hui Zhu	f729b25a10	openapi: Add MemoryConfig balloon Add MemoryConfig balloon to vmm/src/api/openapi/cloud-hypervisor.yaml. Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-07-07 17:25:13 +01:00
Hui Zhu	8b6b97b86f	vmm: Add virtio-balloon support This commit adds new option balloon to memory config. Set it to on will open the balloon function. Signed-off-by: Hui Zhu <teawater@antfin.com>	2020-07-07 17:25:13 +01:00
Rob Bradford	b69f6d4f6c	vhost_user_net, vhost_user_block, option_parser: Remove vmm dependency Remove the vmm dependency from vhost_user_block and vhost_user_net where it was existing to use config::OptionParser. By moving the OptionParser to its own crate at the top-level we can remove the very heavy dependency that these vhost-user backends had. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-07-06 18:33:29 +01:00
Michael Zhao	726e45e0ce	vmm: Divide Seccomp KVM IOCTL rules by architecture Refactored the construction of KVM IOCTL rules for Seccomp. Separating the rules by architecture can reduce the risk of bugs and attacks. Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-06 13:40:38 +01:00
Wei Liu	a4f484bc5e	hypervisor: Define a VM-Exit abstraction In order to move the hypervisor specific parts of the VM exit handling path, we're defining a generic, hypervisor agnostic VM exit enum. This is what the hypervisor's Vcpu run() call should return when the VM exit can not be completely handled through the hypervisor specific bits. For KVM based hypervisors, this means directly forwarding the IO related exits back to the VMM itself. For other hypervisors that e.g. rely on the VMM to decode and emulate instructions, this means the decoding itself would happen in the hypervisor crate exclusively, and the rest of the VM exit handling would be handled through the VMM device model implementation. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com> Fix test_vm unit test by using the new abstraction and dropping some dead code. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-06 12:59:43 +01:00
Wei Liu	cfa758fbb1	vmm, hypervisor: introduce and use make_user_memory_region This removes the last KVM-ism from memory_manager. Also make use of that method in other places. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-06 12:31:19 +02:00
Wei Liu	8d97d628c3	vmm: drop "kvm" from memory slot code The code is purely for maintaining an internal counter. It is not really tied to KVM. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-07-06 12:31:19 +02:00
Samuel Ortiz	8186a8eee6	vmm: interrupt: Rename vm_fd The _fd suffix is KVM specific. But since it now point to an hypervisor agnostic hypervisor::Vm implementation, we should just rename it vm. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Samuel Ortiz	4cc8853fe4	vmm: device_manager: Rename vm_fd The _fd suffix is KVM specific. But since it now point to an hypervisor agnostic hypervisor::Vm implementation, we should just rename it vm. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Samuel Ortiz	2012287611	vmm: memory_manager: Rename fd variable into something more meaningful The fd naming is quite KVM specific. Since we're now using the hypervisor crate abstractions, we can rename those into something more readable and meaningful. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Samuel Ortiz	acfe5eb94f	vmm: vm: Rename fd variable into something more meaningful The fd naming is quite KVM specific. Since we're now using the hypervisor crate abstractions, we can rename those into something more readable and meaningful. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Samuel Ortiz	3db4c003a3	vmm: cpu: Rename fd variable into something more meaningful The fd naming is quite KVM specific. Since we're now using the hypervisor crate abstractions, we can rename those into something more readable and meaningful. Like e.g. vcpu or vm. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Samuel Ortiz	618722cdca	hypervisor: cpu: Rename state getter and setter vcpu.{set_}cpu_state() is a stutter. Signed-off-by: Samuel Ortiz <sameo@linux.intel.com>	2020-07-06 09:35:30 +01:00
Rob Bradford	2a6eb31d5b	vm-virtio, virtio-devices: Split device implementation from virt queues Split the generic virtio code (queues and device type) from the VirtioDevice trait, transport and device implementations. This also simplifies the feature handling in vhost_user_backend as the vm-virtio crate is no longer has any features. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-07-02 17:09:28 +01:00
Michael Zhao	8820e9e133	vmm: Fix Seccomp filter for AArch64 Signed-off-by: Michael Zhao <michael.zhao@arm.com>	2020-07-02 08:46:24 +01:00
Sebastien Boeuf	e35d4c5b28	hypervisor: Store all supported MSRs On x86 architecture, we need to save a list of MSRs as part of the vCPU state. By providing the full list of MSRs supported by KVM, this patch fixes the remaining snapshot/restore issues, as the vCPU is restored with all its previous states. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-30 14:03:03 +01:00
Sebastien Boeuf	e2b5c78dc5	hypervisor: Re-order vCPU state for storing and restoring Some vCPU states such as MP_STATE can be modified while retrieving other states. For this reason, it's important to follow a specific order that will ensure a state won't be modified after it has been saved. Comments about ordering requirements have been copied over from Firecracker commit 57f4c7ca14a31c5536f188cacb669d2cad32b9ca. This patch also set the previously saved VCPU_EVENTS, as this was missing from the restore codepath. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-30 14:03:03 +01:00
Wei Liu	2b8accf49a	vmm: interrupt: put KVM code into a kvm module Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	c31e747005	vmm: interrupt: generify impl InterruptManager for MsiInterruptManager The logic can be shared among hypervisor implementations. The 'static bound is used such that we don't need to deal with extra lifetime parameter everywhere. It should be okay because we know the entry type E doesn't contain any reference. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	ade904e356	vmm: interrupt: generify impl InterruptSourceGroup for MsiInterruptGroup At this point we can use the same logic for all hypervisor implementations. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	2b466ed80c	vmm: interrupt: provide MsiInterruptGroupOps trait Currently it only contains a function named set_gsi_routes. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	b2abead65b	vmm: interrupt: provide and use extension trait RoutingEntryExt This trait contains a function which produces a interrupt routing entry. Implement that trait for KvmRoutingEntry and rewrite the update function. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	4dbca81b86	vmm: interrupt: rename set_kvm_gsi_routes to set_gsi_routes This function will be used to commit routing information to the hypervisor. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	fd7b42e54d	vmm: interrupt: inline mask_kvm_entry The logic for looking up the correct interrupt can be shared among hypervisors. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	0ec39da90c	vmm: interrupt: generify KvmMsiInterruptManager The observation is only the route entry is hypervisor dependent. Keep a definition of KvmMsiInterruptManager to avoid too much code churn. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	d5149e95cb	vmm: interrupt: generify KvmRoutingEntry and KvmMsiInterruptGroup The observation is that only the route field is hypervisor specific. Provide a new function in blanket implementation. Also redefine KvmRoutingEntry with RoutingEntry to avoid code churn. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	637f58bcd9	vmm: interrupt: drop Kvm prefix from KvmLegacyUserspaceInterruptManager This data structure doesn't contain KVM specific stuff. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
Wei Liu	574cab6990	vmm: interrupt: create GSI hashmap directly The observation is that the GSI hashmap remains untouched before getting passed into the MSI interrupt manager. We can create that hashmap directly in the interrupt manager's new function. The drops one import from the interrupt module. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-30 12:09:42 +01:00
dependabot-preview[bot]	f3c8f827cc	build(deps): bump linux-loader from `2a62f21` to `ec930d7` Bumps [linux-loader](https://github.com/rust-vmm/linux-loader) from `2a62f21` to `ec930d7`. - [Release notes](https://github.com/rust-vmm/linux-loader/releases) - [Commits](`2a62f21b44...ec930d700f`) Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com> Signed-off-by: dependabot-preview[bot] <support@dependabot.com>	2020-06-30 07:05:06 +00:00
Rob Bradford	522d8c8412	vmm: openapi: Add the /vm.counters API entry point This is a hash table of string to hash tables of u64s. In JSON these hash tables are object types. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-27 00:07:47 +02:00
Sebastien Boeuf	86377127df	vmm: Resume devices after vCPUs have been resumed Because we don't want the guest to miss any event triggered by the emulation of devices, it is important to resume all vCPUs before we can resume the DeviceManager with all its associated devices. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-25 12:01:34 +02:00
Sebastien Boeuf	f6eeba781b	vmm: Save and restore vCPU states during pause/resume operations We need consistency between pause/resume and snapshot/restore operations. The symmetrical behavior of pausing/snapshotting and restoring/resuming has been introduced recently, and we must now ensure that no matter if we're using pause/resume or snapshot/restore features, the resulting VM should be running in the exact same way. That's why the vCPU state is now stored upon VM pausing. The snapshot operation being a simple serialization of the previously saved state. The same way, the vCPU state is now restored upon VM resuming. The restore operation being a simple deserialization of the previously restored state. It's interesting to note that this patch ensures time consistency from a guest perspective, no matter which clocksource is being used. From a previous patch, the KVM clock was saved/restored upon VM pause/resume. We now have the same behavior for TSC, as the TSC from the vCPUs are saved/restored upon VM pause/resume too. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-25 12:01:34 +02:00
Sebastien Boeuf	18e7d7a1f7	vmm: cpu: Resume before shutdown in a specific way Instead of calling the resume() function from the CpuManager, which involves more than what is needed from the shutdown codepath, and potentially ends up with a deadlock, we replace it with a subset. The full resume operation is reserved for a VM that has been paused. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-25 12:01:34 +02:00
Sebastien Boeuf	65132fb99d	vmm: Implement Pausable trait for Vcpu We want each Vcpu to store the vCPU state upon VM pausing. This is the reason why we need to explicitly implement the Pausable trait for the Vcpu structure. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-25 12:01:34 +02:00
Wei Liu	1741af74ed	hypervisor: add safety statement in set_user_memory_region When set_user_memory_region was moved to hypervisor crate, it was turned into a safe function that wrapped around an unsafe call. All but one call site had the safety statements removed. But safety statement was not moved inside the wrapper function. Add the safety statement back to help reasoning in the future. Also remove that one last instance where the safety statement is not needed . No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-25 10:25:13 +02:00
Wei Liu	b27439b6ed	arch, hypervisor, vmm: KvmHyperVisor -> KvmHypervisor "Hypervisor" is one word. The "v" shouldn't be capitalised. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-25 10:25:13 +02:00
Wei Liu	b00171e17d	vmm: use MemoryRegion where applicable That removes one more KVM-ism in VMM crate. Note that there are more KVM specific code in those files to be split out, but we're not at that stage yet. No functional change. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-25 10:25:13 +02:00
Rob Bradford	d983c0a680	vmm: Expose counters from virtio devices to API Collate the virtio device counters in DeviceManager for each device that exposes any and expose it through the recently added HTTP API. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-25 07:02:44 +02:00
Rob Bradford	bca8a19244	vmm: Implement HTTP API for obtaining counters The counters are a hash of device name to hash of counter name to u64 value. Currently the API is only implemented with a stub that returns an empty set of counters. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-25 07:02:44 +02:00
Rob Bradford	fd4aba8eae	vmm: api: Implement support for GET handlers EndpointHandler This can be used for simple API requests which return data but do not require any input. Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-25 07:02:44 +02:00
Rob Bradford	80be393b16	vmm: api: Order HTTP entry points in alphabetical order Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-25 07:02:44 +02:00
Wei Liu	4cc37d7b9a	vmm: interrupt: drop a few pub keywords Those items are not used elsewhere. Restrict their scope. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-24 12:39:42 +02:00
Wei Liu	1661adbbaf	vmm: interrupt: add "Kvm" prefix to MsiInterruptGroup The structure is tightly coupled with KVM. It uses KVM specific structures and calls. Add Kvm prefix to it. Microsoft hypervisor will implement its own interrupt group(s) later. No functional change intended. Signed-off-by: Wei Liu <liuwe@microsoft.com>	2020-06-24 12:39:42 +02:00
Sebastien Boeuf	9f4714c32a	vmm: Extend seccomp filters with KVM_KVMCLOCK_CTRL Now that the VMM uses KVM_KVMCLOCK_CTRL from the KVM API, it must be added to the seccomp filters list. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-24 12:38:56 +02:00
Sebastien Boeuf	4a81d65f79	vmm: Notify the guest about vCPUs being paused Through the newly added API notify_guest_clock_paused(), this patch improves the vCPU pause operation by letting the guest know that each vCPU is being paused. This is important to avoid soft lockups detection from the guest that could happen because the VM has been paused for more than 20 seconds. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-24 12:38:56 +02:00
Sebastien Boeuf	9fa8438063	vmm: Fill CpuManager's vCPU list on restore path It's important that on restore path, the CpuManager's vCPU gets filled with each new vCPU that is being created. In order to cover both boot and restore paths, the list is being filled from the common function create_vcpu(). Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-24 12:38:56 +02:00
Sebastien Boeuf	f5150aa261	vmm: Extend seccomp filters with KVM_GET_CLOCK and KVM_SET_CLOCK Now that the VMM uses both KVM_GET_CLOCK and KVM_SET_CLOCK from the KVM API, they must be added to the seccomp filters list. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-23 14:36:01 +01:00
Sebastien Boeuf	8038161861	vmm: Get and set clock during pause and resume operations In order to maintain correct time when doing pause/resume and snapshot/restore operations, this patch stores the clock value on pause, and restore it on resume. Because snapshot/restore expects a VM to be paused before the snapshot and paused after the restore, this covers the migration use case too. Signed-off-by: Sebastien Boeuf <sebastien.boeuf@intel.com>	2020-06-23 14:36:01 +01:00
Rob Bradford	4b64f2a027	vmm: cpu: Reuse already allocated vCPUs if available When a request is made to increase the number of vCPUs in the VM attempt to reuse any previously removed (and hence inactive) vCPUs before creating new ones. This ensures that the APIC ID is not reused for a different KVM vCPU (which is not allowed) and that the APIC IDs are also sequential. The two key changes to support this are: * Clearing the "kill" bit on the old vCPU state so that it does not immediately exit upon thread recreation. * Using the length of the vcpus vector (the number of allocated vcpus) rather than the number of active vCPUs (.present_vcpus()) to determine how many should be created. This change also introduced some new info!() debugging on the vCPU creation/removal path to aid further development in the future. TEST=Expanded test_cpu_hotplug test. Fixes: #1338 Signed-off-by: Rob Bradford <robert.bradford@intel.com>	2020-06-23 14:11:14 +01:00

... 3 4 5 6 7 ...

1221 Commits