This document covers the command line options which the Xen Hypervisor.
Most parameters take the form option=value. Different
options on the command line should be space delimited. All options are
case sensitive, as are all values unless explicitly noted.
<boolean>)All boolean option may be explicitly enabled using a
value of > yes, on,
true, enable or 1
They may be explicitly disabled using a value of >
no, off, false,
disable or 0
In addition, a boolean option may be enabled by simply stating its
name, and may be disabled by prefixing its name with
no-.
####Examples
Enable noreboot mode > noreboot=true
Disable x2apic support (if present) > x2apic=off
Enable synchronous console mode > sync_console
Explicitly specifying any value other than those listed above is
undefined, as is stacking a no- prefix with an explicit
value.
<integer>)An integer parameter will default to decimal and may be prefixed with
a - for negative numbers. Alternatively, a hexadecimal
number may be used by prefixing the number with 0x, or an
octal number may be used if a leading 0 is present.
Providing a string which does not validly convert to an integer is undefined.
<size>)A size parameter may be any integer, with a single size suffix
T or t: TiB (2^40)G or g: GiB (2^30)M or m: MiB (2^20)K or k: KiB (2^10)B or b: BytesWithout a size suffix, the default will be kilo. Providing a suffix other than those listed above is undefined.
Many parameters are more complicated and require more intricate configuration. The detailed description of each individual parameter specify which values are valid.
Some options take a comma separated list of values.
Some parameters act as combinations of the above, most commonly a mix of Boolean and String. These are noted in the relevant sections.
= force | ht | noirq | <boolean> | verbose
String, or Boolean to disable.
By default, Xen will scan the DMI data and blacklist certain systems
which are known to have broken ACPI setups. Providing
acpi=force will cause Xen to ignore the blacklist and
attempt to use all ACPI features.
Using acpi=ht causes Xen to parse the ACPI tables enough
to enumerate all CPUs, but will not use other ACPI features. This is not
common, and only has an effect if your system is blacklisted.
The acpi=noirq option causes Xen to not parse the ACPI
MADT table looking for IO-APIC entries. This is also not common, and any
system which requires this option to function should be blacklisted.
Additionally, this will not prevent Xen from finding IO-APIC entries
from the MP tables.
Further, any of the boolean false options can be used to disable ACPI usage entirely.
Because responsibility for ACPI processing is shared between Xen and the domain 0 kernel this option is automatically propagated to the domain 0 command line.
Finally, acpi=verbose will enable per-processor
information logging which may otherwise be too noisy in particular on
large systems.
= <integer>
Specify which ACPI MADT table to parse for APIC information, if more than one is present.
= <boolean>
Default:
false
Enforce checking that P-state transitions by the ACPI cpufreq driver actually result in the nominated frequency to be established. A warning message will be logged if that isn’t the case.
= <boolean>
Instruct Xen to ignore timer-interrupt override.
= s3_bios | s3_mode
s3_bios instructs Xen to invoke video BIOS
initialization during S3 resume.
s3_mode instructs Xen to set up the boot time (option
vga=) video mode during S3 resume.
= <boolean>
Default:
false
Force boot on potentially unsafe systems. By default Xen will refuse to boot on systems with the following errata:
= <boolean>
Default:
false
Permit multiple copies of host p2m.
= bigsmp | default
Override Xen’s logic for choosing the APIC driver. By default, if
there are more than 8 CPUs, Xen will switch to bigsmp over
default.
= <boolean>
Default:
true
Permit Xen to use APIC Virtualisation Extensions. This is an optimisation available as part of VT-x, and allows hardware to take care of the guests APIC handling, rather than requiring emulation in Xen.
= verbose | debug
Increase the verbosity of the APIC code from the default value.
= <boolean>
Default:
true
Permit Xen to use “Always Running APIC Timer” support on compatible hardware in combination with cpuidle. This option is only expected to be useful for developers wishing Xen to fall back to older timing methods on newer hardware.
= List of [ <bool>, mac-permissive=<bool> ]
Controls for the Argo hypervisor-mediated interdomain communication service.
The functionality that this option controls is only available when Xen has been compiled with the build setting for Argo enabled in the build configuration.
Argo is a interdomain communication mechanism, where Xen acts as the central point of authority. Guests may register memory rings to recieve messages, query the status of other domains, and send messages by hypercall, all subject to appropriate auditing by Xen. Argo is disabled by default.
The mac-permissive boolean controls whether wildcard
receive rings may be registered (mac-permissive=1) or may
not be registered (mac-permissive=0).
This option is disabled by default, to protect domains from a DoS by a buggy or malicious other domain spamming the ring.
= <boolean>
Default:
true
Permit Xen to use Address Space Identifiers. This is an optimisation which tags the TLB entries with an ID per vcpu. This allows for guest TLB flushes to be performed without the overhead of a complete TLB flush.
= <boolean>
Default:
false
Forces all CPUs’ full state to be logged upon certain fatal asynchronous exceptions (watchdog NMIs and unexpected MCEs).
= <boolean>
Default:
false
Permits Xen to set up and use PCI Address Translation Services. This is a performance optimisation for PCI Passthrough.
WARNING: Xen cannot currently safely use ATS because of its synchronous wait loops for Queued Invalidation completions.
= <size>
Default:
0(no limit)
Specify a maximum amount of available memory, to which Xen will clamp the e820 table.
= List of [ <integer> | <integer>-<integer> ]
Specify that certain pages, or certain ranges of pages contain bad
bytes and should not be used. For example, if your memory tester says
that byte 0x12345678 is bad, you would place
badpage=0x12345 on Xen’s command line.
= idle | <boolean>
Default:
idle
Scrub free RAM during boot. This is a safety feature to prevent accidentally leaking sensitive VM data into other VMs if Xen crashes and reboots.
In idle mode, RAM is scrubbed in background on all CPUs
during idle-loop with a guarantee that memory allocations always provide
scrubbed pages. This option reduces boot time on machines with a large
amount of RAM while still providing security benefits.
= <size>
Default:
128M
Maximum RAM block size chunks to be scrubbed whilst holding the page heap lock and not running softirqs. Reduce this if softirqs are not being run frequently enough. Setting this to a high value may cause boot failure, particularly if the NMI watchdog is also enabled.
= List of [ shstk=<bool>, ibt=<bool> ]
Applicability: x86
Controls for the use of Control-flow Enforcement Technology. CET is group a of hardware features designed to combat Return-oriented Programming (ROP, also call/jmp COP/JOP) attacks.
CET is incompatible with 32bit PV guests. If any CET sub-options are
active, they will override the pv=32 boolean to
false. Backwards compatibility can be maintained with the
pv-shim mechanism.
The shstk= boolean controls whether Xen uses Shadow
Stacks for its own protection.
The option is available when CONFIG_XEN_SHSTK is
compiled in, and generally defaults to true on hardware
supporting CET-SS. Specifying cet=no-shstk will cause Xen
not to use Shadow Stacks even when support is available in hardware.
Some hardware suffers from an issue known as Supervisor Shadow Stack
Fracturing. On such hardware, Xen will default to not using Shadow
Stacks when virtualised. Specifying cet=shstk will override
this heuristic and enable Shadow Stacks unilaterally.
The ibt= boolean controls whether Xen uses Indirect
Branch Tracking for its own protection.
The option is available when CONFIG_XEN_IBT is compiled
in, and defaults to true on hardware supporting CET-IBT.
Specifying cet=no-ibt will cause Xen not to use Indirect
Branch Tracking even when support is available in hardware.
= pit | hpet | acpi | tsc
If set, override Xen’s default choice for the platform timer. Having TSC as platform timer requires being explicitly set. This is because TSC can only be safely used if CPU hotplug isn’t performed on the system. On some platforms, the “maxcpus” option may need to be used to further adjust the number of allowed CPUs. When running on platforms that can guarantee a monotonic TSC across sockets you may want to adjust the “tsc” command line parameter to “stable:socket”.
= <integer>
Default:
2
Specify the event count threshold for raising Corrected Machine Check Interrupts. Specifying zero disables CMCI handling.
= <boolean>
Default:
false
Flag to indicate whether to probe for a CMOS Real Time Clock irrespective of ACPI indicating none to be there.
= <baud>[/<base-baud>][,[DPS][,[<io-base>|pci|amt][,[<irq>|msi][,[<port-bdf>][,[<bridge-bdf>]]]]]]
Both option com1 and com2 follow the same
format.
<baud> may be either an integer baud rate, or the
string auto if the bootloader or other earlier firmware has
already set it up.DPS represents the number of data bits, the parity, and
the number of stop bits.
D is an integer between 5 and 8 for the number of data
bits.P is a single character representing the type of
parity:
n Noo Odde Evenm Marks SpaceS is an integer 1 or 2 for the number of stop
bits.<io-base> is an integer which specifies the IO
base port for UART registers.<irq> is the IRQ number to use, or 0
to use the UART in poll mode only, or msi to set up a
Message Signaled Interrupt.<port-bdf> is the PCI location of the UART, in
<bus>:<device>.<function> notation.<bridge-bdf> is the PCI bridge behind which is
the UART, in <bus>:<device>.<function>
notation.pci indicates that Xen should scan the PCI bus for the
UART, avoiding Intel AMT devices.amt indicated that Xen should scan the PCI bus for the
UART, including Intel AMT devices if present.A typical setup for most situations might be
com1=115200,8n1
In addition to the above positional specification for UART parameters, name=value pair specfications are also supported. This is used to add flexibility for UART devices which require additional UART parameter configurations.
The comma separation still delineates positional parameters. Hence, unless the parameter is explicitly specified with name=value option, it will be considered a positional parameter.
The syntax consists of com1=(comma-separated positional parameters),(comma separated name-value pairs)
The accepted name keywords for name=value pairs are:
baud - accepts integer baud rate (eg. 115200) or
autobridge- Similar to bridge-bdf in positional parameters.
Used to determine the PCI bridge to access the UART device. Notation is
xx:xx.x <bus>:<device>.<function>clock-hz- accepts large integers to setup UART clock
frequencies. Do note - these values are multiplied by 16.data-bits - integer between 5 and 8dev - accepted values are pci OR
amt. If this option is used to specify if the serial device
is pci-based. The io_base cannot be specified when dev=pci
or dev=amt is used.io-base - accepts integer which specified IO base port
for UART registersirq - IRQ number to useparity - accepted values are same as positional
parametersport - Used to specify which port the PCI serial device
is located on Notation is xx:xx.x
<bus>:<device>.<function>reg-shift - register shifts required to set UART
registersreg-width - register width required to set UART
registers (only accepts 1 and 4)stop-bits - only accepts 1 or 2 for the number of stop
bitsThe following are examples of correct specifications:
com1=115200,8n1,0x3f8,4
com1=115200,8n1,0x3f8,4,reg-width=4,reg-shift=2
com1=baud=115200,parity=n,stop-bits=1,io-base=0x3f8,reg-width=4
= <size>
Default:
conring_size=16k
Specify the size of the console ring buffer.
= List of [ vga | com1[H,L] | com2[H,L] | pv | dbgp | ehci | xhci | none ]
Default:
console=com1,vga
Specify which console(s) Xen should use.
vga indicates that Xen should try and use the vga
graphics adapter.
com1 and com2 indicates that Xen should use
serial ports 1 and 2 respectively. Optionally, these arguments may be
followed by an H or L. H
indicates that transmitted characters will have their MSB set, while
received characters must have their MSB set. L indicates
the converse; transmitted and received characters will have their MSB
cleared. This allows a single port to be shared by two subsystems
(e.g. console and debugger).
pv indicates that Xen should use Xen’s PV console. This
option is only available when used together with
pv-in-pvh.
dbgp or ehci indicates that Xen should use
a USB2 debug port.
xhci indicates that Xen should use a USB3 debug
port.
none indicates that Xen should not use a console. This
option only makes sense on its own.
= none | date | datems | boot | raw
Default:
none
Can be modified at runtime
Specify which timestamp format Xen should use for each console line.
none: No timestampsdate: Date and time information
[YYYY-MM-DD HH:MM:SS]datems: Date and time, with milliseconds
[YYYY-MM-DD HH:MM:SS.mmm]boot: Seconds and microseconds since boot
[SSSSSS.uuuuuu]raw: Raw platform ticks, architecture and
implementation dependent
[XXXXXXXXXXXXXXXX]For compatibility with the older boolean parameter, specifying
console_timestamps alone will enable the date
option.
= <boolean>
Default:
false
Flag to indicate whether all guest console output should be copied into the console ring buffer.
= <switch char>[x]
Default:
conswitch=a
Can be modified at runtime
Specify which character should be used to switch serial input between Xen and dom0. The required sequence is CTRL-<switch char> three times.
The optional trailing x indicates that Xen should not
automatically switch the console input to dom0 during boot. Any other
value, including omission, causes Xen to automatically switch to the
dom0 console during dom0 boot. Use conswitch=ax to keep the
default switch character, but for xen to keep the console.
= power | performance
Default:
power
= arch_perfmon
If set, force use of the performance counters for oprofile, rather than detecting available support.
= none | {{ <boolean> | xen } [:[powersave|performance|ondemand|userspace][,<maxfreq>][,[<minfreq>][,[verbose]]]]} | dom0-kernel
Default:
xen
Indicate where the responsibility for driving power states lies. Note
that the choice of dom0-kernel is deprecated and not
supported by all Dom0 kernels.
<maxfreq> and <minfreq> are
integers which represent max and min processor frequencies
respectively.verbose option can be included as a string or also as
verbose=<integer>
= List of comma separated booleans
This option allows for fine tuning of the facilities Xen will use, after accounting for hardware capabilities as enumerated via CPUID.
Unless otherwise noted, options only have any effect in their negative form, to hide the named feature(s). Ignoring a feature using this mechanism will cause Xen not to use the feature, nor offer them as usable to guests.
Currently accepted:
The Speculation Control hardware features srbds-ctrl,
md-clear, ibrsb, stibp,
ibpb, l1d-flush and ssbd are used
by default if available and applicable. They can all be ignored.
rdrand and rdseed have multiple
interactions.
For Special Register Buffer Data Sampling (SRBDS, XSA-320, CVE-2020-0543), RDRAND and RDSEED can be ignored.
Due to the absence of microcode to address SRBDS on IvyBridge client
hardware, the RDRAND feature is hidden by default for guests, unless
rdrand is used in its positive form. Irrespective of the
setting here, VMs can use RDRAND if explicitly enabled in guest config
file, and VMs already using RDRAND can migrate in.
The RDRAND feature is disabled by default on AMD Fam15/16
systems, due to possible malfunctions after ACPI S3 suspend/resume.
rdrand may be used in its positive form to override Xen’s
default behaviour on these systems, and make the feature fully
usable.
= fam_0f_rev_[cdefg] | fam_10_rev_[bc] | fam_11_rev_b
Applicability: AMD
If none of the other cpuid_mask_* options are given, Xen has a set of pre-configured masks to make the current processor appear to be family/revision specified.
See below for general information on masking.
Warning: This option is not fully effective on Family 15h processors or later.
= <integer>
Applicability: x86. Default:
~0(all bits set)
The availability of these options are model specific. Some processors don’t support any of them, and no processor supports all of them. Xen will ignore options on processors which are lacking support.
These options can be used to alter the features visible via the
CPUID instruction. Settings applied here take effect
globally, including for Xen and all guests.
Note: Since Xen 4.7, it is no longer necessary to mask a host to create migration safety in heterogeneous scenarios. All necessary CPUID settings should be provided in the VM configuration file. Furthermore, it is recommended not to use this option, as doing so causes an unnecessary reduction of features at Xen’s disposal to manage guests.
= <boolean>
= <boolean>
= <string>
Can be modified at runtime
Specify debug-key actions in cases of crashes. Each of the parameters
applies to a different crash reason. The <string> is
a sequence of debug key characters, with + having the
special meaning of a 10 millisecond pause.
crash-debug-debugkey will be used for crashes induced by
the C debug key (i.e. manually induced crash).
crash-debug-hwdom denotes a crash of dom0.
crash-debug-kexeccmd is an explicit request of dom0 to
continue with the kdump kernel via kexec. Only available on hypervisors
built with CONFIG_KEXEC.
crash-debug-panic is a crash of the hypervisor.
crash-debug-watchdog is a crash due to the watchdog
timer expiring.
It should be noted that dumping diagnosis data to the console can fail in multiple ways (missing data, hanging system, …) depending on the reason of the crash, which might have left the hypervisor in a bad state. In case a debug-key action leads to another crash recursion will be avoided, so no additional debug-key actions will be performed in this case. A crash in the early boot phase will not result in any debug-key action, as the system might not yet be in a state where the handlers can work.
So e.g. crash-debug-watchdog=0+0r would dump dom0 state
twice with 10 milliseconds between the two state dumps, followed by the
run queues of the hypervisor, if the system crashes due to a watchdog
timeout.
Depending on the reason of the system crash it might happen that triggering some debug key action will result in a hang instead of dumping data and then doing a reboot or crash dump.
= <size>
Default:
4G
Specify the maximum address to allocate certain structures, if used in combination with the low_crashinfo command line option.
= <ramsize-range>:<size>[,...][{@,<}<offset>]= <size>[{@,<}<offset>]= <size>,below=offset
Specify sizes and optionally placement of the crash kernel
reservation area. The <ramsize-range>:<size>
pairs indicate how much memory to set aside for a crash kernel
(<size>) for a given range of installed RAM
(<ramsize-range>). Each
<ramsize-range> is of the form
<start>-[<end>].
A trailing @<offset> specifies the exact address
this area should be placed at, whereas < in place of
@ just specifies an upper bound of the address range the
area should fall into.
< and below are synonyomous, the latter being useful for grub2 systems which would otherwise require escaping of the < option
= <integer>
= <integer>
= <integer>
Default:
10
Domains subject to a cap receive a replenishment of their runtime budget once every cap period interval. Default is 10 ms. The amount of budget they receive depends on their cap. For instance, a domain with a 50% cap will receive 50% of 10 ms, so 5 ms.
= <integer>
Default:
18
Specify the number of bits to use for the fractional part of the values involved in Credit2 load tracking and load balancing math.
= <integer>
Default:
30
Specify the number of bits to use to represent the length of the window (in nanoseconds) we use for load tracking inside Credit2. This means that, with the default value (30), we use 2^30 nsec ~= 1 sec long window.
Load tracking is done by means of a variation of exponentially weighted moving average (EWMA). The window length defined here is what tells for how long we give value to previous history of the load itself. In fact, after a full window has passed, what happens is that we discard all previous history entirely.
A short window will make the load balancer quick at reacting to load changes, but also short-sighted about previous history (and hence, e.g., long term load trends). A long window will make the load balancer thoughtful of previous history (and hence capable of capturing, e.g., long term load trends), but also slow in responding to load changes.
The default value of 1 sec is rather long.
= cpu | core | socket | node | all
Default:
socket
Specify how host CPUs are arranged in runqueues. Runqueues are kept
balanced with respect to the load generated by the vCPUs running on
them. Smaller runqueues (as in with core) means more
accurate load balancing (for instance, it will deal better with
hyperthreading), but also more overhead.
Available alternatives, with their meaning, are: * cpu:
one runqueue per each logical pCPUs of the host; * core:
one runqueue per each physical core of the host; * socket:
one runqueue per each physical socket (which often, but not always,
matches a NUMA node) of the host; * node: one runqueue per
each NUMA node of the host; * all: just one runqueue shared
by all the logical pCPUs of the host
Regardless of the above choice, Xen attempts to respect
sched_credit2_max_cpus_runqueue limit, which may mean more
than one runqueue for the all value. If that isn’t
intended, raise the sched_credit2_max_cpus_runqueue
value.
= ehci[ <integer> | @pci<bus>:<slot>.<func> ]= xhci[ <integer> | @pci<bus>:<slot>.<func> ][,share=<bool>|hwdom]
Specify the USB controller to use, either by instance number (when going over the PCI busses sequentially) or by PCI device (must be on segment 0).
Use ehci for EHCI debug port, use xhci for
XHCI debug capability. XHCI driver will wait indefinitely for the debug
host to connect - make sure the cable is connected. The
share option for xhci controls who else can use the
controller: * no: use the controller exclusively for
console, even hardware domain (dom0) cannot use it * hwdom:
hardware domain may use the controller too, ports not used for debug
console will be available for normal devices; this is the default *
yes: the controller can be assigned to any domain; it is
not safe to assign the controller to untrusted domain
Choosing share=hwdom (the default) or
share=yes allows a domain to reset the controller, which
may cause small portion of the console output to be lost.
The share=yes configuration is not security
supported.
= <integer>
Default:
20
Limits the number lines printed in Xen stack traces.
= [cpu:]<size>
Default:
128
Specify the size of the console debug trace buffer. By specifying
cpu: additionally a trace buffer of the specified size is
allocated per cpu. The debug trace feature is only enabled in debugging
builds of Xen.
= <boolean>
Default:
CONFIG_DIT_DEFAULT
Specify whether Xen and guests should operate in Data Independent Timing mode (Intel calls this DOITM, Data Operand Independent Timing Mode). Note that enabling this option cannot guarantee anything beyond what underlying hardware guarantees (with, where available and known to Xen, respective tweaks applied).
= <integer>
Specify the bit width of the DMA heap.
= List of [ pv | pvh, shadow=<bool>, verbose=<bool>,
cpuid-faulting=<bool>, msr-relaxed=<bool> ]
Applicability: x86
Controls for how dom0 is constructed on x86 systems.
The pv and pvh options select the
virtualisation mode of dom0.
The pv option is only available when
CONFIG_PV is compiled in. The pvh option is
only available when CONFIG_HVM is compiled in. When both
options are compiled in, the default is PV.
In addition, the following requirements must be met:
The shadow boolean allows dom0 to be explicitly
constructed using shadow paging. This option is unavailable when
CONFIG_SHADOW_PAGING is disabled.
For PVH, dom0 defaults to using HAP on capable hardware, and falls back to shadow paging otherwise. A PVH dom0 cannot be used if Xen is compiled without shadow paging support, and the hardware lacks HAP support.
For PV, the use of dom0 shadow mode is only for development purposes. PV guests do no require any paging support by default.
The verbose boolean is intended for diagnostics, and
prints out extra information during the dom0 build. It defaults to the
compile time choice of CONFIG_VERBOSE_DEBUG.
The cpuid-faulting boolean is an interim option, is
only applicable to PV dom0, and defaults to true.
Before Xen 4.13, the domain builder logic for guest construction depended on seeing host CPUID values to function correctly. As a result, CPUID Faulting was never activated for PV dom0’s, even on capable hardware.
In Xen 4.13, the domain builder logic has been fixed, and no longer has this dependency. As a consequence, CPUID Faulting is activated by default even for PV dom0’s.
However, as PV dom0’s have always seen host CPUID data in the past,
there is a chance that further dependencies exist. This boolean can be
used to restore the pre-4.13 behaviour. If specifying
no-cpuid-faulting fixes an issue in dom0, please report a
bug.
The msr-relaxed boolean is an interim option, and
defaults to false.
In Xen 4.15, the default behaviour for unhandled MSRs has been changed, to avoid leaking host data into guests, and to avoid breaking guest logic which uses #GP probing to identify the availability of MSRs.
However, this new stricter behaviour has the possibility to break
guests, and a more 4.14-like behaviour can be selected by specifying
dom0=msr-relaxed.
If using this option is necessary to fix an issue, please report a bug.
= List of comma separated booleans
Applicability: x86
This option allows for fine tuning of the facilities dom0 will use, after accounting for hardware capabilities and Xen settings as enumerated via CPUID.
Options are accepted in positive and negative form, to enable or disable specific features. All selections via this mechanism are subject to normal CPU Policy safety and dependency logic.
This option is intended for developers to opt dom0 into non-default features, and is not intended for use in production circumstances. If using this option is necessary to fix an issue, please report a bug.
= List of [ passthrough=<bool>, strict=<bool>, map-inclusive=<bool>,
map-reserved=<bool>, none ]
Controls for the dom0 IOMMU setup.
The passthrough boolean controls whether IOMMU
translation functionality is disabled for devices in dom0
(passthrough=1) or whether the IOMMU is used to ensure that
dom0 can only DMA to its permitted areas of RAM
(passthrough=0).
This option is only applicable to x86 PV dom0’s, and defaults to false.
Some older Intel VT-d hardware isn’t capable of disabling translation functionality on a per-device basis, and will cause this option to be ignored and assumed to be 0. Similar behaviour on such systems is only available by fully disabling all IOMMUs.
This option is hardwired to false for x86 PVH dom0’s (where a non-identity transform is required for dom0 to function), and is ignored for ARM.
The strict boolean is applicable to x86 PV dom0’s
only and defaults to false. It controls whether dom0 can have IOMMU
mappings for all domain RAM in the system, or only for its allocated RAM
(and grant mappings etc.)
This option is hardwired to true for x86 PVH dom0’s (as RAM belonging to other domains in the system don’t live in a compatible address space), and is ignored for ARM.
The map-inclusive boolean is applicable to x86 PV
dom0’s, and sets up identity IOMMU mappings for all non-RAM regions
below 4GB except for unusable ranges, and ranges belonging to Xen.
Typically, some devices in a system use bits of RAM for communication, and these areas should be listed as reserved in the E820 table and identified via RMRR or IVMD entries in the ACPI tables, so Xen can ensure that they are identity-mapped in the IOMMU. However, some firmware makes mistakes, and this option is a coarse-grain workaround for those errors.
Where possible, finer grain corrections should be made with the
rmrr=, ivmd=, ivrs_hpet[]=, or
ivrs_ioapic[]= command line options.
This option is disabled by default, and deprecated and intended for
removal in future versions of Xen. If specifying
map-inclusive is the only way to make your system boot,
please report a bug.
The map-reserved functionality is very similar to
map-inclusive.
The differences from map-inclusive are that
map-reserved is applicable to both x86 PV and PVH dom0’s,
is enabled by default, and represents a subset of the correction by only
mapping reserved memory regions rather than all non-RAM
regions.
The none option is intended for development purposes
only, and skips certain safety checks pertaining to the correct IOMMU
configuration for dom0 to boot.
Incorrect use of this option may result in a malfunctioning system.
= List of <hex>-<hex>
Specify a list of IO ports to be excluded from dom0 access.
Either:
= <integer>.
The number of VCPUs to give to dom0. This number of VCPUs can be more than the number of PCPUs on the host. The default is the number of PCPUs.
Or:
= <min>-<max>where<min>and<max>are integers.
Gives dom0 a number of VCPUs equal to the number of PCPUs, but always
at least <min> and no more than
<max>. Using <min> may give more
VCPUs than PCPUs. <min> or <max>
may be omitted and the defaults of 1 and unlimited respectively are used
instead.
For example, with dom0_max_vcpus=4-8:
Number of PCPUs | Dom0 VCPUs 2 | 4 4 | 4 6 | 6 8 | 8 10 | 8
= <size>
Set the amount of memory for the initial domain (dom0). It must be greater than zero. This parameter is required.
= List of ( min:<sz> | max:<sz> | <sz> )
Set the amount of memory for the initial domain (dom0). If a size is positive, it represents an absolute value. If a size is negative, it is subtracted from the total available memory.
<sz> specifies the exact amount of memory.min:<sz> specifies the minimum amount of
memory.max:<sz> specifies the maximum amount of
memory.If <sz> is not specified, the default is all the
available memory minus some reserve. The reserve is 1/16 of the
available memory or 128 MB (whichever is smaller).
The amount of memory will be at least the minimum but never more than
the maximum (i.e., max overrides the min
option). If there isn’t enough memory then as much as possible is
allocated.
max:<sz> also sets the maximum reservation (the
maximum amount of memory dom0 can balloon up to). If this is omitted
then the maximum reservation is unlimited.
For example, to set dom0’s initial memory allocation to 512MB but
allow it to balloon up as far as 1GB use
dom0_mem=512M,max:1G
<sz>is:<size> | [<size>+]<frac>%<frac>is an integer < 100
<frac> specifies a fraction of host memory size
in percent.So <sz> being 1G+25% on a 256 GB host
would result in 65 GB.
If you use this option then it is highly recommended that you disable any dom0 autoballooning feature present in your toolstack. See the xl.conf(5) man page or Xen Best Practices.
This option doesn’t have effect if pv-shim mode is enabled.
= List of [ <integer> | relaxed | strict ]
Default:
strict
Specify the NUMA nodes to place Dom0 on. Defaults for vCPU-s created
and memory assigned to Dom0 will be adjusted to match the node
restrictions set up here. Note that the values to be specified here are
ACPI PXM ones, not Xen internal node numbers. relaxed sets
up vCPU affinities to prefer but be not limited to the specified
node(s).
= <boolean>
Default:
false
Pin dom0 vcpus to their respective pcpus
= path [:options]
Default:
""
Specify the full path in the device tree for the UART. If the path
doesn’t start with /, it is assumed to be an alias. The
options are device specific.
= <boolean>
Flag that specifies if RAM should be clipped to the highest cacheable MTRR.
Default:
trueon Intel CPUs, otherwisefalse
= <boolean>
Default:
false
Flag that enables verbose output when processing e820 information and applying clipping.
= off | on | skipmbr
Control retrieval of Extended Disc Data (EDD) from the BIOS during boot.
= no | force
Either force retrieval of monitor EDID information via VESA DDC, or disable it (edid=no). This option should not normally be required except for debugging purposes.
= List of [ rs=<bool>, attr=no|uc ]
Controls for interacting with the system Extended Firmware Interface.
The rs boolean controls whether Runtime Services are
used. By default, Xen uses Runtime Services itself, and proxies certain
calls on behalf of dom0. Selecting rs=0 prohibits all use
of Runtime Services.
The attr= string exists to specify what to do with
memory regions of unknown/unrecognised cacheability.
attr=no is the default and will leave the memory regions
unmapped, while attr=uc will map them as fully
uncacheable.
= List of [ ad=<bool>, pml=<bool>, exec-sp=<bool> ]
Applicability: Intel
Extended Page Tables are a feature of Intel’s VT-x technology, whereby hardware manages the virtualisation of HVM guest pagetables. EPT was introduced with the Nehalem architecture.
The ad boolean controls hardware tracking of Access
and Dirty bits in the EPT pagetables, and was first introduced in
Broadwell Server.
By default, Xen will use A/D tracking when available in hardware,
except on Avoton processors affected by erratum AVR41. Explicitly
choosing ad=0 will disable the use of A/D tracking on
capable hardware, whereas choosing ad=1 will cause tracking
to be used even on AVR41-affected hardware.
The pml boolean controls the use of Page
Modification Logging, which is also introduced in Broadwell Server.
PML is a feature whereby the processor generates a list of pages which have been dirtied. This is necessary information for operations such as live migration, and having the processor maintain the list of dirtied pages is more efficient than traditional software implementations where all guest writes trap into Xen so the dirty bitmap can be maintained.
By default, Xen will use PML when it is available in hardware. PML
functionally depends on A/D tracking, so choosing ad=0 will
implicitly disable PML. pml=0 can be used to prevent the
use of PML on otherwise capable hardware.
The exec-sp boolean controls whether EPT superpages
with execute permissions are permitted. In general this is good for
performance.
However, on processors vulnerable CVE-2018-12207, HVM guest kernels can use executable superpages to crash the host. By default, executable superpages are disabled on affected hardware.
If HVM guest kernels are trusted not to mount a DoS against the system, this option can enabled to regain performance.
This boolean may be modified at runtime using
xl set-parameters ept=[no-]exec-sp to switch between fast
and secure.
When switching from secure to fast, preexisting HVM domains will run at their current performance until they are rebooted; new domains will run without any overhead.
When switching from fast to secure, all HVM domains will immediately suffer a performance penalty.
Warning: No guarantee is made that this runtime option will be retained indefinitely, or that it will retain this exact behaviour. It is intended as an emergency option for people who first chose fast, then change their minds to secure, and wish not to reboot.
= [<domU number>][,<dom0 number>]
Default:
32,<variable>
Change the number of PIRQs available for guests. The optional first
number is common for all domUs, while the optional second number
(preceded by a comma) is for dom0. Changing the setting for domU has no
impact on dom0 and vice versa. For example to change dom0 without
changing domU, use extra_guest_irqs=,512. The default value
for Dom0 and an eventual separate hardware domain is architecture
dependent. The upper limit for both values on x86 is such that the
resulting total number of IRQs can’t be higher than 32768. Note that
specifying zero as domU value means zero, while for dom0 it means to use
the default.
= <boolean>
Default :
true
Flag to enable or disable support for extended regions for Dom0 and Dom0less DomUs.
Extended regions are ranges of unused address space exposed to the guest as “safe to use” for special memory mappings. Disable if your board device tree is incomplete.
= permissive | enforcing | late | disabled
Default:
enforcing
Specify how the FLASK security server should be configured. This option is only available if the hypervisor was compiled with FLASK support. This can be enabled by running either: - make -C xen config and enabling XSM and FLASK. - make -C xen menuconfig and enabling ‘FLux Advanced Security Kernel support’ and ‘Xen Security Modules support’
permissive: This is intended for development and is not
suitable for use with untrusted guests. If a policy is provided by the
bootloader, it will be loaded; errors will be reported to the ring
buffer but will not prevent booting. The policy can be changed to
enforcing mode using “xl setenforce”.enforcing: This will cause the security server to enter
enforcing mode prior to the creation of domain 0. If an valid policy is
not provided by the bootloader and no built-in policy is present, the
hypervisor will not continue booting.late: This disables loading of the built-in security
policy or the policy provided by the bootloader. FLASK will be enabled
but will not enforce access controls until a policy is loaded by a
domain using “xl loadpolicy”. Once a policy is loaded, FLASK will run in
enforcing mode unless “xl setenforce” has changed that setting.disabled: This causes the XSM framework to revert to
the dummy module. The dummy module provides the same security policy as
is used when compiling the hypervisor without support for XSM. The
xsm_op hypercall can also be used to switch to this mode after boot, but
there is no way to re-enable FLASK once the dummy module is loaded.
= <height>where height is8x8 | 8x14 | 8x16
Specify the font size when using the VESA console driver.
= <boolean>
Default:
false
Allow EPT to be enabled when VMX feature
VM_ENTRY_LOAD_GUEST_PAT is not present.
Warning: Due to CVE-2013-2212, VMX feature
VM_ENTRY_LOAD_GUEST_PAT is by default required as a
prerequisite for using EPT. If you are not using PCI Passthrough, or
trust the guest administrator who would be using passthrough, then the
requirement can be relaxed. This option is particularly useful for
nested virtualization, to allow the L1 hypervisor to use EPT even if the
L0 hypervisor does not provide VM_ENTRY_LOAD_GUEST_PAT.
= com1[H,L] | com2[H,L] | dbgp
Default: ``
Specify which console gdbstub should use. See console.
= List of [ max-ver:<integer>, transitive=<bool>, transfer=<bool> ]
Default (Arm):
gnttab=max-ver:1Default (x86,PV):gnttab=max-ver:2,transitive,transferDefault (x86,HVM):gnttab=max-ver:2,transitive
Control various aspects of the grant table behaviour available to guests.
max-ver Select the maximum grant table version to offer
to guests. Valid version are 1 and 2.transitive Permit or disallow the use of transitive
grants. Note that the use of grant table v2 without transitive grants is
an ABI breakage from the guests point of view.transfer Permit or disallow the GNTTABOP_transfer
operation of the grant table hypercall. Note that disallowing
GNTTABOP_transfer is an ABI breakage from the guests point of view. This
option is only available on hypervisors configured to support PV
guests.The usage of gnttab v2 is not security supported on ARM platforms.
= <integer>
Default:
64
Can be modified at runtime
Specify the maximum number of frames which any domain may use as part of its grant table. This value is an upper boundary of the per-domain value settable via Xen tools.
Dom0 is using this value for sizing its grant table.
= <integer>
Default:
1024
Can be modified at runtime
Specify the maximum number of frames to use as part of a domains maptrack array. This value is an upper boundary of the per-domain value settable via Xen tools.
= <boolean>
Applicability: x86
Default: true unless running virtualized on AMD or Hygon hardware
Control whether to use global pages for PV guests, and thus the need to perform TLB flushes by writing to CR4. This is a performance trade-off.
AMD SVM does not support selective trapping of CR4 writes, which means that a global TLB flush (two CR4 writes) takes two VMExits, and massively outweigh the benefit of using global pages to begin with. This case is easy for Xen to spot, and is accounted for in the default setting.
Other cases where this option might be a benefit is on VT-x hardware when selective CR4 writes are not supported/enabled by the hypervisor, or in any virtualised case using shadow paging. These are not easy for Xen to spot, so are not accounted for in the default setting.
= <level>[/<rate-limited level>]where level isnone | error | warning | info | debug | all
Default:
guest_loglvl=none/warning
Can be modified at runtime
Set the logging level for Xen guests. Any log message with equal more more importance will be printed.
The optional <rate-limited level> option instructs
which severities should be rate limited.
= <boolean>
Default:
true
Flag to globally enable or disable support for Hardware Assisted Paging (HAP)
= <boolean>
Default:
true
Flag to enable 1 GB host page table support for Hardware Assisted Paging (HAP).
= <boolean>
Default:
true
Flag to enable 2 MB host page table support for Hardware Assisted Paging (HAP).
= <domid>
Default:
0
Enable late hardware domain creation using the specified domain ID. This is intended to be used when domain 0 is a stub domain which builds a disaggregated system including a hardware domain with the specified domain ID. This option is supported only when compiled with XSM on x86.
= <boolean>
Default:
false
Control Xens use of the APEI Hardware Error Source Table, should one be found.
= <size>
Specify the memory boundary past which memory will be treated as highmem (x86 debug hypervisor only).
= <boolean>
Default :
false
Say yes at your own risk if you want to enable heterogenous computing (such as big.LITTLE). This may result to an unstable and insecure platform, unless you manually specify the cpu affinity of all domains so that all vcpus are scheduled on the same class of pcpus (big or LITTLE but not both). vcpu migration between big cores and LITTLE cores is not supported. See docs/misc/arm/big.LITTLE.txt for more information.
When the hmp-unsafe option is disabled (default), CPUs that are not identical to the boot CPU will be parked and not used by Xen.
= List of [ <bool> | broadcast=<bool> | legacy-replacement=<bool> ]
Applicability: x86
Controls Xen’s use of the system’s High Precision Event Timer. By
default, Xen will use an HPET when available and not subject to errata.
Use of the HPET can be disabled by specifying hpet=0.
The broadcast boolean is disabled by default, but
forces Xen to keep using the broadcast for CPUs in deep C-states even
when an RTC interrupt is enabled. This then also affects raising of the
RTC interrupt.
The legacy-replacement boolean allows for control
over whether Legacy Replacement mode is enabled.
Legacy Replacement mode is intended for hardware which does not have an 8254 PIT, and allows the HPET to be configured into a compatible mode. Intel chipsets from Skylake/ApolloLake onwards can turn the PIT off for power saving reasons, and there is no platform-agnostic mechanism for discovering this.
By default, Xen will not change hardware configuration, unless the PIT appears to be absent, at which point Xen will try to enable Legacy Replacement mode before falling back to pre-IO-APIC interrupt routing options.
This behaviour can be inhibited by specifying
legacy-replacement=0. Alternatively, this mode can be
enabled unconditionally (if available) by specifying
legacy-replacement=1.
= <boolean>
Deprecated alternative of hpet=broadcast.
= <integer>
The specified value is a bit mask with the individual bits having the following meaning:
Bit 0 - debug level 0 (unused at present) Bit 1 - debug level 1 (Control Register logging) Bit 2 - debug level 2 (VMX logging of MSR restores when context switching) Bit 3 - debug level 3 (unused at present) Bit 4 - I/O operation logging Bit 5 - vMMU logging Bit 6 - vLAPIC general logging Bit 7 - vLAPIC timer logging Bit 8 - vLAPIC interrupt logging Bit 9 - vIOAPIC logging Bit 10 - hypercall logging Bit 11 - MSR operation logging
Recognized in debug builds of the hypervisor only.
= <boolean>
Default:
false
Allow use of the Forced Emulation Prefix in HVM guests, to allow emulation of arbitrary instructions.
This option is intended for development and testing purposes.
Warning As this feature opens up the instruction emulator to arbitrary instruction from an HVM guest, don’t use this in production system. No security support is provided when this flag is set.
= <boolean>
Default:
true
Specify whether guests are to be given access to physical port 80 (often used for debugging purposes), to override the DMI based detection of systems known to misbehave upon accesses to that port.
= <integer>
= old | new
Default:
newunless directed-EOI is supported
= List of [ <bool>, verbose, debug, force, required,
quarantine=<bool>|scratch-page,
sharept, superpages, intremap, intpost, crash-disable,
snoop, qinval, igfx, amd-iommu-perdev-intremap,
dom0-{passthrough,strict} ]
All sub-options are boolean in nature.
I/O Memory Memory Units perform a function similar to the CPU MMU (hence the name), but typically exist as a discrete device, integrated as part of a PCI Root Complex. The most common configuration is to have one IOMMU per package (for on-die PCIe devices and directly attached PCIe lanes), and one IOMMU covering the remaining I/O in the system.
The functionality in an IOMMU commonly falls into two orthogonal categories:
DMA remapping which uses a pagetable-like hierarchical structure and maps I/O Virtual Addresses (DFNs - Device Frame Numbers in Xen’s terminology) to System Physical Addresses (MFNs - Machine Frame Numbers in Xen’s terminology).
Interrupt Remapping, which controls incoming Message Signalled Interrupt requests, including their routing to specific CPUs.
IOMMU functionality ca