Skip to content

Document that host-switch isn't needed for shared-storage vMotion - #9

Open
doccaz wants to merge 5 commits into
cloudbase:masterfrom
doccaz:host-switch-investigation
Open

doccaz wants to merge 5 commits into
cloudbase:masterfrom
doccaz:host-switch-investigation

Conversation

@doccaz

@doccaz doccaz commented Sep 20, 2026 •

Copy link
Copy Markdown

Summary

Investigated the "Host-switch AIO messages" (NFC_AIO_SWITCH_HOST_*) backlog item. The binary's own strings (SWITCHHOST_VADP, a PreSwitchHost callback string carrying a full new-host descriptor) show this is VDDK's mechanism for keeping an NFC/backup session alive across a live vMotion. Testing it needed a second ESXi host in the cluster with shared storage.

Three findings, fully closing the question:

  1. Compute-only vMotion (shared NFS datastore) while an SSL-hook-captured native-VDDK NFC read session stayed open — completely unaffected.
  2. Combined storage+compute vMotion — a disk on host-local VMFS the target host had no path to at all, relocated to shared storage and the compute moved to that host, in one RelocateVM_Task. Still completely unaffected, including reads issued after the migration fully completed. Both times the wire capture shows a single TCP fd for the entire session, no reconnect, no NFC_AIO_SWITCH_HOST_* traffic.
  3. No public API surface for the mechanism at all — grepped every header in the VDDK 8.0.3 SDK for SwitchHost/Callback and found only the documented, unrelated completion/progress/logging callbacks. There is no public registration function for PreSwitchHost anywhere. Whatever this mechanism does, it's wired into VMware's own internal/first-party tooling (VADP), not reachable through any API a third-party client — or OpenVixDiskLib — actually links against. This closes the investigation without needing to test the remaining scenario (the connected host itself becoming unavailable) by disrupting a real lab host — there's no code path by which a normal client could ever need to implement this regardless.
  • docs/host_switch.md — the investigation and evidence (both vMotion tests plus the API-surface finding)
  • docs/host_switch_lab_setup.md — the 2-host/shared-storage lab build (reused an existing NFS server already running in the underlying cluster rather than standing up new storage)
  • docs/reverse_engineering_procedure.md / README.md updated

Also found, and corrected mid-PR: an earlier commit on this branch mischaracterized a real finding as NFS-specific — that opening a running VM's disk over NFC needs a snapshot. It doesn't matter which datastore type; confirmed by powering on SLES16 (an ordinary long-lived VMFS-backed lab VM) and reproducing the identical failure. This project's own lab pytest fixture never powers on its temp VM, so every prior capture in this whole project had been reading a powered-off VM's disk without that being a deliberate choice.

Test plan

  • Compute-only vMotion with an active NFC session — zero effect
  • Combined storage+compute vMotion with an active NFC session — zero effect, verified via wire capture (single fd, no reconnect)
  • Confirmed the "needs a snapshot when powered on" finding is general, not NFS-specific, against a second VM on VMFS
  • Confirmed no public VDDK SDK API exists for the host-switch callback mechanism (all three headers checked)

Note on PR sequencing

Stacked on top of the unmerged direct-ESXi/GetInfo/QueryAllocatedBlocks/CBT work (PRs #3–#6), same as #7 and #8 — this branch's history includes their combined commit, so the diff shown here will shrink to just the three new commits (a84a0ae, 133457e, bf99794) once those merge. Docs-only, no code changes, no overlap with the other PRs.

Reverse-engineered and implemented against a live standalone ESXi 8.0.3
host (no vCenter), closing most of the gap versus the proprietary VDDK:

- Direct ESXi (no vCenter) connectivity: nfc_service() previously
  hardcoded the NfcService moref as "nfcService" (vCenter's name),
  which fails on bare ESXi (moref is "ha-nfc-service" there). Now
  resolved dynamically via RetrieveInternalContent, same as VDDK
  itself does. Also fixes connect_authd() for tickets that omit
  `host` (implicit on a direct-ESXi ticket).

- VixDiskLib_GetInfo: capacity and physical geometry come free from
  the OPEN_FILE reply (offsets already in the wire frame). biosGeo,
  adapterType, and uuid are fetched via DDB_GET, matching real VDDK's
  behavior and cost exactly.

- DDB_GET (VMDK descriptor lookups): generic key/value NFC message,
  values are ASCII text on the wire (not binary), matching how a VMDK
  descriptor's DDB section is stored.

- VixDiskLib_QueryAllocatedBlocks: allocated-block bitmap query.
  Verified against a live disk to exactly match native VDDK's output,
  including two non-obvious wire details: a field-order swap that's
  invisible in a zero-offset capture, and 4-byte bitmap padding that
  only shows up for small chunk counts.

- Changed Block Tracking: turned out to need no NFC work at all --
  VirtualMachine.QueryChangedDiskAreas is public VIM API. Added thin
  wrappers (enable_change_tracking / disk_change_id /
  query_changed_disk_areas) and documented real-world characteristics
  (extent granularity, wildcard changeId semantics) from live testing.

- Investigated NFC_DELTA_DISK: found it's an optional VMFS-only VDDK
  client optimization (per `strings` on libvixDiskLib.so), not a
  correctness requirement -- reading, writing, and querying allocated
  blocks on an actual snapshot delta file already work with the
  existing NFC_DISK-only implementation. Documented a real gotcha
  found along the way: querying allocated blocks on the same
  still-open handle a write just went through can see stale data.

Adds unit tests (bitmap decode/merge, DDB_GET wire format, CBT
dataclass conversion, validation errors) and integration tests
(GetInfo, QueryAllocatedBlocks, CBT full cycle, delta-disk read/write/
query) validated against a live ESXi 8.0.3 lab. Full protocol details
and the reverse-engineering process are in docs/nfc_auth.md,
docs/nfc_open.md, docs/nfc_read.md, docs/cbt.md, and
docs/reverse_engineering_procedure.md.
Investigated the "Host-switch AIO messages" (NFC_AIO_SWITCH_HOST_*)
backlog item, which the binary's own strings show is VDDK's mechanism
for keeping an NFC/backup session alive across a live vMotion. Testing
it needed a second ESXi host in the cluster with shared storage -- a
real infra build, documented separately.

With two hosts sharing an NFS datastore, kept a native-VDDK NFC read
session alive under the SSL hook while triggering a live vMotion
mid-session. Result: completely unaffected -- a single TCP file
descriptor served the whole session, no reconnect, no SWITCH_HOST
traffic at all. NFC access is datastore-based, not VM/host-based, so a
compute-only vMotion with shared storage never needs anything to
change on the NFC side. Documented what this doesn't rule out (Storage
vMotion, losing the connected host) rather than treating it as a
closed question.

Also recorded a real gotcha hit while building the test case: NFC
needs a snapshot to open a *running* VM's disk on an NFS datastore
(works immediately on VMFS, which every other capture in this project
has used).
@doccaz
doccaz marked this pull request as ready for review September 20, 2026 14:03
…n test

The previous commit's "NFC needs a snapshot on NFS" framing was wrong.
Confirmed directly: opening a running (powered-on) VM's disk over NFC
fails identically on VMFS -- tested against SLES16, an ordinary
long-lived lab VM that had simply never been powered on during any
prior capture in this project, because the lab pytest fixture's temp
VM never is either. The real constraint is "powered on without a
snapshot", full stop, regardless of datastore type.

Also ran the more aggressive test this correction motivated: a combined
storage+compute vMotion, relocating a VM's disk from host-local VMFS
(unreachable by the target host) to shared NFS while migrating compute
to that same target host in one call. Still no effect on an open NFC
session -- single TCP fd for the whole migration, no reconnect, no
NFC_AIO_SWITCH_HOST_* traffic, even after the relocation completed.
Narrows what's left untested down to one case: the connected host
itself becoming unavailable, independent of any migration.
The connected-host-unavailable case (the last untested scenario)
turned out not to need testing: grepped every header in the VDDK 8.0.3
SDK for SwitchHost/Callback and found no public registration function
for the PreSwitchHost mechanism the binary strings reference -- only
the documented, unrelated completion/progress/logging callbacks.

Whatever NFC_AIO_SWITCH_HOST_*/PreSwitchHost actually do, they're
wired into VMware's own internal/first-party tooling, not exposed
through any API a third-party backup vendor -- or OpenVixDiskLib --
actually links against. There is no code path by which a normal
VixDiskLib_Open/Read/Write client could ever trigger, observe, or need
to implement this, regardless of what happens to the underlying hosts.
This closes the investigation with no remaining open questions.
@doccaz
doccaz marked this pull request as ready for review September 20, 2026 14:54
Comment thread docs/host_switch_lab_setup.md Outdated
`ks=` argument in directly (no live interaction needed at boot at
all), or just budget for one interactive install.

**Also learned**: opening a VM's console **twice** (e.g. clicking

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This document contains some potentially useful information but I'm not sure if we should document the vCenter / vMotion deployment procedure here, especially referencing scripts that aren't included in the repo.

Drop environment-specific deployment details and references to
scripts that are not part of the repo, per review feedback.
@doccaz

doccaz commented Sep 23, 2026

Copy link
Copy Markdown
Author

Good point, thanks. I trimmed docs/host_switch_lab_setup.md down to the generic requirements for a 2-host vMotion lab: shared datastore, second host in the same cluster, vMotion enabled on a VMkernel adapter, and the RelocateVM_Task check. I removed the environment-specific parts (how the NFS server and the nested host were provisioned) and every reference to scripts that aren't in the repo.

Those references pointed at a local .lab/ directory that we never committed. It holds the scripts I used for the encryption and host-switch investigations:

  • Lab setup: vc_setup.py and move_host_to_cluster.py (add a host to vCenter and move it into a cluster), mount_nfs_datastore.py (mount the same NFS export on both hosts), esx_capacity.py and esx_netinfo.py (host RAM and network checks), and a VCSA deploy config template.
  • Encryption: nkp_setup.py (creates a Native Key Provider over REST), create_encrypted_vm.py, encrypt_with_vtpm.py, encrypt_disk.py, check_disk_encryption.py, and relocate_encrypted_vm_cross_host.py (cold-relocates an encrypted VM to a host that has never held its key).
  • Host-switch tests: trigger_vmotion.py and trigger_storage_vmotion.py trigger a live or storage vMotion while vddk_loop_read_trace.py holds an NFC read session open under an SSL hook. Both tests were negative: one TCP fd for the whole session, and no NFC_AIO_SWITCH_HOST_* traffic.
  • Tracing: sslhook.c and vddk_open_crypto_trace.py (VDDK tracing through an SSL hook), and ovdl_crypto_readwrite.py (a read/write test on an encrypted disk).

They currently have hard-coded assumptions about my lab and would need sanitizing first. Would you like me to add any of them to the repo? If not, the docs stand on their own now.

@petrutlucian94

Copy link
Copy Markdown
Member

Thanks for cleaning up the the docs. I'd keep it simple for now so that we can merge it more easily.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants