Conversation
Reverse-engineered and implemented against a live standalone ESXi 8.0.3 host (no vCenter), closing most of the gap versus the proprietary VDDK: - Direct ESXi (no vCenter) connectivity: nfc_service() previously hardcoded the NfcService moref as "nfcService" (vCenter's name), which fails on bare ESXi (moref is "ha-nfc-service" there). Now resolved dynamically via RetrieveInternalContent, same as VDDK itself does. Also fixes connect_authd() for tickets that omit `host` (implicit on a direct-ESXi ticket). - VixDiskLib_GetInfo: capacity and physical geometry come free from the OPEN_FILE reply (offsets already in the wire frame). biosGeo, adapterType, and uuid are fetched via DDB_GET, matching real VDDK's behavior and cost exactly. - DDB_GET (VMDK descriptor lookups): generic key/value NFC message, values are ASCII text on the wire (not binary), matching how a VMDK descriptor's DDB section is stored. - VixDiskLib_QueryAllocatedBlocks: allocated-block bitmap query. Verified against a live disk to exactly match native VDDK's output, including two non-obvious wire details: a field-order swap that's invisible in a zero-offset capture, and 4-byte bitmap padding that only shows up for small chunk counts. - Changed Block Tracking: turned out to need no NFC work at all -- VirtualMachine.QueryChangedDiskAreas is public VIM API. Added thin wrappers (enable_change_tracking / disk_change_id / query_changed_disk_areas) and documented real-world characteristics (extent granularity, wildcard changeId semantics) from live testing. - Investigated NFC_DELTA_DISK: found it's an optional VMFS-only VDDK client optimization (per `strings` on libvixDiskLib.so), not a correctness requirement -- reading, writing, and querying allocated blocks on an actual snapshot delta file already work with the existing NFC_DISK-only implementation. Documented a real gotcha found along the way: querying allocated blocks on the same still-open handle a write just went through can see stale data. Adds unit tests (bitmap decode/merge, DDB_GET wire format, CBT dataclass conversion, validation errors) and integration tests (GetInfo, QueryAllocatedBlocks, CBT full cycle, delta-disk read/write/ query) validated against a live ESXi 8.0.3 lab. Full protocol details and the reverse-engineering process are in docs/nfc_auth.md, docs/nfc_open.md, docs/nfc_read.md, docs/cbt.md, and docs/reverse_engineering_procedure.md.
Investigated the "Host-switch AIO messages" (NFC_AIO_SWITCH_HOST_*) backlog item, which the binary's own strings show is VDDK's mechanism for keeping an NFC/backup session alive across a live vMotion. Testing it needed a second ESXi host in the cluster with shared storage -- a real infra build, documented separately. With two hosts sharing an NFS datastore, kept a native-VDDK NFC read session alive under the SSL hook while triggering a live vMotion mid-session. Result: completely unaffected -- a single TCP file descriptor served the whole session, no reconnect, no SWITCH_HOST traffic at all. NFC access is datastore-based, not VM/host-based, so a compute-only vMotion with shared storage never needs anything to change on the NFC side. Documented what this doesn't rule out (Storage vMotion, losing the connected host) rather than treating it as a closed question. Also recorded a real gotcha hit while building the test case: NFC needs a snapshot to open a *running* VM's disk on an NFS datastore (works immediately on VMFS, which every other capture in this project has used).
…n test The previous commit's "NFC needs a snapshot on NFS" framing was wrong. Confirmed directly: opening a running (powered-on) VM's disk over NFC fails identically on VMFS -- tested against SLES16, an ordinary long-lived lab VM that had simply never been powered on during any prior capture in this project, because the lab pytest fixture's temp VM never is either. The real constraint is "powered on without a snapshot", full stop, regardless of datastore type. Also ran the more aggressive test this correction motivated: a combined storage+compute vMotion, relocating a VM's disk from host-local VMFS (unreachable by the target host) to shared NFS while migrating compute to that same target host in one call. Still no effect on an open NFC session -- single TCP fd for the whole migration, no reconnect, no NFC_AIO_SWITCH_HOST_* traffic, even after the relocation completed. Narrows what's left untested down to one case: the connected host itself becoming unavailable, independent of any migration.
The connected-host-unavailable case (the last untested scenario) turned out not to need testing: grepped every header in the VDDK 8.0.3 SDK for SwitchHost/Callback and found no public registration function for the PreSwitchHost mechanism the binary strings reference -- only the documented, unrelated completion/progress/logging callbacks. Whatever NFC_AIO_SWITCH_HOST_*/PreSwitchHost actually do, they're wired into VMware's own internal/first-party tooling, not exposed through any API a third-party backup vendor -- or OpenVixDiskLib -- actually links against. There is no code path by which a normal VixDiskLib_Open/Read/Write client could ever trigger, observe, or need to implement this, regardless of what happens to the underlying hosts. This closes the investigation with no remaining open questions.
| `ks=` argument in directly (no live interaction needed at boot at | ||
| all), or just budget for one interactive install. | ||
|
|
||
| **Also learned**: opening a VM's console **twice** (e.g. clicking |
There was a problem hiding this comment.
This document contains some potentially useful information but I'm not sure if we should document the vCenter / vMotion deployment procedure here, especially referencing scripts that aren't included in the repo.
Drop environment-specific deployment details and references to scripts that are not part of the repo, per review feedback.
|
Good point, thanks. I trimmed Those references pointed at a local
They currently have hard-coded assumptions about my lab and would need sanitizing first. Would you like me to add any of them to the repo? If not, the docs stand on their own now. |
|
Thanks for cleaning up the the docs. I'd keep it simple for now so that we can merge it more easily. |
Summary
Investigated the "Host-switch AIO messages" (
NFC_AIO_SWITCH_HOST_*) backlog item. The binary's own strings (SWITCHHOST_VADP, aPreSwitchHost callbackstring carrying a full new-host descriptor) show this is VDDK's mechanism for keeping an NFC/backup session alive across a live vMotion. Testing it needed a second ESXi host in the cluster with shared storage.Three findings, fully closing the question:
RelocateVM_Task. Still completely unaffected, including reads issued after the migration fully completed. Both times the wire capture shows a single TCP fd for the entire session, no reconnect, noNFC_AIO_SWITCH_HOST_*traffic.SwitchHost/Callbackand found only the documented, unrelated completion/progress/logging callbacks. There is no public registration function forPreSwitchHostanywhere. Whatever this mechanism does, it's wired into VMware's own internal/first-party tooling (VADP), not reachable through any API a third-party client — or OpenVixDiskLib — actually links against. This closes the investigation without needing to test the remaining scenario (the connected host itself becoming unavailable) by disrupting a real lab host — there's no code path by which a normal client could ever need to implement this regardless.docs/host_switch.md— the investigation and evidence (both vMotion tests plus the API-surface finding)docs/host_switch_lab_setup.md— the 2-host/shared-storage lab build (reused an existing NFS server already running in the underlying cluster rather than standing up new storage)docs/reverse_engineering_procedure.md/README.mdupdatedAlso found, and corrected mid-PR: an earlier commit on this branch mischaracterized a real finding as NFS-specific — that opening a running VM's disk over NFC needs a snapshot. It doesn't matter which datastore type; confirmed by powering on
SLES16(an ordinary long-lived VMFS-backed lab VM) and reproducing the identical failure. This project's ownlabpytest fixture never powers on its temp VM, so every prior capture in this whole project had been reading a powered-off VM's disk without that being a deliberate choice.Test plan
Note on PR sequencing
Stacked on top of the unmerged direct-ESXi/GetInfo/QueryAllocatedBlocks/CBT work (PRs #3–#6), same as #7 and #8 — this branch's history includes their combined commit, so the diff shown here will shrink to just the three new commits (
a84a0ae,133457e,bf99794) once those merge. Docs-only, no code changes, no overlap with the other PRs.