Commit Graph

22 Commits

Author SHA1 Message Date
Kevin Adams 415c2a250b fix: unmap targetextent before extent delete in free_image (#259)
When a VM has multiple disks on the same per-VM target, migrating one
disk while the VM is running leaves the target in use for the other
disks. TrueNAS refuses to delete an extent associated with an in-use
target even with force=true.

Fix: delete the targetextent (LUN mapping) first, then delete the
extent. Removing the mapping severs this disk's association without
touching the active session or other LUNs on the same target.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:37:33 -04:00
Kevin Adams 5e2e5a8cbe fix: allow free_image on detached disks while VM is running (#258)
Removed _vm_is_running() check from free_image. When a disk is detached
from a running VM (unused0), QEMU immediately closes its libiscsi
connection — there is no active session to block the delete. The
force=true flag on the TrueNAS extent DELETE is the correct guard for
any residual session. The VM-running check was overly broad and blocked
legitimate deletes of already-detached disks.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:28:32 -04:00
Kevin Adams d7b0735902 refactor: per-VM iSCSI targets with iscsi:// paths (#255)
Each VM now gets its own TrueNAS iSCSI target (proxmox-vm-<vmid>)
instead of sharing a single storage-wide target.  path() returns
iscsi://portal/iqn:proxmox-vm-<vmid>/lun — QEMU connects via libiscsi,
the same pattern ZFSPlugin.pm uses.

When a VM stops, QEMU closes its libiscsi connection.  TrueNAS sees
no active session on that VM's target, so extent DELETE with force=true
succeeds without stopping the iSCSI service.  Deleting one VM's disk
while other VMs are running on the same storage now works correctly.

Changes:
- Remove all iscsiadm session management (_iscsi_ensure_session,
  _iscsi_session_exists, _wait_for_device, _dev_path,
  _running_vms_on_storage, _resolve_target)
- Add _resolve_vm_target: find-or-create proxmox-vm-<vmid> target,
  inherit portal/initiator groups from existing targets
- Add _maybe_cleanup_vm_target: delete empty VM target after last disk
- Add _vm_is_running: check owning VM before deletion (uses PID file)
- Add _api_global: cached iSCSI global config (basename)
- path(): returns iscsi://portal/per-vm-iqn/lun (target looked up by
  target_id from extent, so legacy shared-target disks still work)
- activate_storage: API reachability check only
- activate_volume: verify extent exists, QEMU handles the connection
- deactivate_storage/volume: clear cache, return 1
- free_image: check owning VM stopped, force-DELETE extent, zvol,
  cleanup empty target — no service-stop fallback needed

Closes #255

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 15:04:01 -04:00
Kevin Adams 01318919fe fix: untaint PID in _running_vms_on_storage for PVE taint-mode Perl
PVE runs pvedaemon with -T (taint mode). Data read from files is tainted
and cannot be passed to kill() without explicit untainting. Extract PID
via regex capture which Perl treats as safe.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:35:46 -04:00
Kevin Adams 5684b335a1 fix: pre-flight check before service-stop fallback in free_image (#253)
Stopping the TrueNAS iSCSI service to clear session state (required on
CORE 13 when force=true is ignored) disrupts all LUNs on the target,
causing io-error on any other running VM using this storage.

Before triggering the service-stop path, scan /etc/pve/qemu-server/*.conf
for running VMs (PID file exists + process alive) that reference this
storage ID. If any are found, fail with a clear message naming the
blocking VMs rather than silently disrupting them.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:29:14 -04:00
Kevin Adams 178088197f fix: service-stop fallback for free_image on CORE 13 (#251)
force=true on extent/targetextent DELETE is unconditionally ignored by
CORE 13.0 — it always checks for active sessions. POST /service/restart
is async and iscsid reconnects in milliseconds, so the restart window
was never clear enough for the DELETE to succeed.

New fallback sequence (only triggered when force=true fails):
1. Set iscsiadm node to manual startup — prevents auto-reconnect
2. iscsiadm --logout — drops initiator session
3. POST /service/stop + poll until service is stopped (up to 15s)
4. DELETE targetextent — succeeds with service fully stopped
5. DELETE extent
6. POST /service/start — restore service
7. Restore automatic startup + discovery + login + rescan

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:20:11 -04:00
Kevin Adams c1f021c5a6 fix: hybrid free_image fallback — service restart + explicit targetextent delete (#251)
force=true on extent DELETE is not sufficient when CORE 13 has an active
or recovering session holding the targetextent (422 persists through retries).

Fallback path (triggered only when force fails and a targetextent exists):
1. POST /service/restart — purges all server-side session state
2. DELETE /iscsi/targetextent/id/{id} — now succeeds with service down
3. DELETE /iscsi/extent/id/{id} — clean delete with no association

Fast path (force=true) is tried first and handles the normal case without
any service disruption.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:24:36 -04:00
Kevin Adams 9cd3cfa234 refactor: remove initiator logout/login from free_image
Session management was only needed to prevent 422 when explicitly
deleting the targetextent.  Now that we delete the extent first with
force=true (which cascades targetextent removal server-side), the
initiator session can stay up throughout — same as the original
LunCmd/FreeNAS.pm delete path which never touched iscsiadm.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:09:55 -04:00
Kevin Adams 377503f89d fix: follow original FreeNAS.pm delete order — extent with force cascades targetextent
The original LunCmd/FreeNAS.pm (v2.0 path) deleted the extent with
{ force: true } and skipped the explicit targetextent DELETE entirely,
relying on TrueNAS to cascade it.  Our v3.0 had the order backwards:
deleting the targetextent first always hit 422 "target in use" when
any iSCSI session was active.

New approach mirrors the original:
1. Log out initiator session
2. DELETE extent with { force: true } — TrueNAS cascades targetextent
3. If that fails (CORE 13 strict session enforcement), restart iSCSI
   service and retry the extent DELETE (up to 5 times)
4. Restore session

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:03:38 -04:00
Kevin Adams 594b533a06 fix: volume_size_info override + free_image retry loop + remove rootdir content type
volume_size_info: base class routes through filesystem_path() → get_subdir()
which dies on block storage. Override to query TrueNAS API directly for
volsize.parsed so create_efidisk and other callers get the correct zvol size.

free_image retry loop: TrueNAS CORE 13 refuses targetextent deletion while any
iSCSI session is active. Single logout+restart was racy — pvedaemon workers
reconnect between the restart and DELETE when multiple disks are deleted
concurrently (e.g. VM destroy). Retry up to 5 times with fresh logout+service
restart each cycle and exponential backoff.

Remove rootdir from plugindata content types: rootdir signals LXC
container/directory storage and caused PVE to route TPM state allocation to
TrueNAS, which immediately failed since TrueNAS has no filesystem path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 08:57:25 -04:00
Kevin Adams f521089e57 fix: parse_volname + free_image session/LUN teardown for CORE 13
parse_volname: implement for vm-/base- volume names so PVE can resolve
device paths and hotplug disks. Without this the base class falls through
to directory-volume parsing and fails with a 400 hotplug error.

free_image: TrueNAS CORE 13 holds iSCSI sessions in recovery state after
TCP disconnect, blocking targetextent deletion with 422 even seconds after
logout. Fix: log out the initiator session by SID, then restart the TrueNAS
iSCSI service to immediately clear server-side session state. Delete the
targetextent and extent, then restore the initiator session so other LUNs
on the same target remain accessible.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 08:12:41 -04:00
Kevin Adams c6f662b88c fix: rescan iSCSI session in _wait_for_device to discover new LUNs
After TrueNAS reloads its iSCSI service, the PVE host's existing session
does not automatically pick up new LUNs. Adding iscsiadm --rescan every 5s
during the device wait loop lets the kernel discover newly exported LUNs
without requiring a full session logout/login.

Fixes the 30s hotplug timeout seen when adding a disk to a running VM.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 22:28:26 -04:00
Kevin Adams 835648fd69 fix: use /pool/dataset for space stats; default shared=1 for iSCSI
TrueNAS CORE 13.0 /pool API does not expose top-level size/free/allocated
fields (they are nested in topology.data[].stats). Switch status() to query
/pool/dataset?id=<pool> which has available.parsed + used.parsed on both
CORE and SCALE.

Add shared=1 as the default in the UI panel — iSCSI is a network block
device accessible from all cluster nodes, so it should be shared storage
by default.

Verified on pve01-hq against Tank01 (CORE 13.0-U6.7):
  TrueNAS01-Tank01  active  1804599296  165936464  1638662832  9.20%

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 22:17:18 -04:00
Kevin Adams 9c85fdb6e9 fix: extract pool name from path in status(); improve Pool/Dataset UX
status() was comparing the full dataset path (e.g. tank/proxmox/vdisks)
against TrueNAS /pool names which are top-level only (e.g. tank).
Extract the first path component so pool stats resolve correctly.

Rename Pool field label to 'Pool / Dataset Path' and update its hint
to make clear it accepts the full ZFS path (matching v2.x 'pool' field).
Rename Dataset to 'Sub-dataset' with a hint that discourages filling it
unless you genuinely need an extra sub-level beyond what's in Pool.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:52:41 -04:00
Kevin Adams 132875a5c5 fix: replace /iscsi/targetgroup with target.groups[] for CORE compatibility
TrueNAS CORE 13.0 does not expose /iscsi/targetgroup as a REST endpoint
(returns 404). Both CORE and SCALE include the portal group associations
inline in each target's 'groups' array from GET /iscsi/target, which is
all we need to filter targets by reachable portal IP.

Remove the /iscsi/targetgroup call entirely — use target.groups[].portal
cross-referenced against /iscsi/portal results instead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:42:19 -04:00
Kevin Adams c5868dd4f6 fix: remove sensitive-properties — store API key in storage.cfg
PVE's sensitive-properties mechanism extracts listed keys from $param
before check_config and passes them only to on_add_hook/on_update_hook.
activate_storage reads from $cfg which never receives those values, so
the API key was always missing at runtime.

The API key now lives in storage.cfg (root-readable, mode 0640, same as
the v2.x truenas_secret field). Proper on_add_hook private-file storage
is tracked in issue #247.

Restore truenas_api_key => {} (required on create). PVE's update flow
passes $create=0 to check_config which skips absent keys, so
edit-without-changing-key still works.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:34:36 -04:00
Kevin Adams efcad405e5 fix: sensitive-property handling + UI autofill + API key reveal button
truenas_api_key must be optional in options() because PVE extracts
sensitive-properties from the POST body before calling check_config,
causing a spurious 'missing required option' 500 on storage create.
Add a guard in _api() so a missing key produces a clear error.

UI: add autocomplete="url" on host and "new-password" on API key to
prevent Firefox/Chrome from filling in PVE login credentials.
Add a reveal trigger button to show/hide the API key field.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:17:58 -04:00
Kevin Adams 8293a718a3 ci: rewrite build pipeline for v3.0 — no patches, TrueNAS.pm + JS only
Remove the validate-patches job (no more patch files in v3.0).
Update lint to target TrueNAS.pm instead of FreeNAS.pm.
Rebuild staging assembly: TrueNAS.pm + truenas-storage.js are the
only payload — no patch dirs, no REST-Client.pm, no triggers file.
Add $VERSION = '3.0.0' to TrueNAS.pm as the single source of truth.
Add release/3.x as a testing-channel branch alongside master.
Pin actions to @v4 (upload/download-artifact, checkout).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:57:18 -04:00
Kevin Adams f6ca0c9ef5 fix: correct package name and API version for PVE 8.4.x
Package must match filename (TrueNAS.pm → PVE::Storage::Custom::TrueNAS)
so PVE's module loader finds the class when checking ISA PVE::Storage::Plugin.

api() returns 11 to match APIVER in PVE::Storage on PVE 8.4.x.
Returning 10 loaded successfully but triggered a deprecation warning.

Verified: pvedaemon loads plugin cleanly with no errors or warnings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:36:13 -04:00
Kevin Adams 432be1ba5c feat: implement TrueNAS v3.0 custom storage plugin (#225, #226, #227)
Full PVE::Storage::Custom implementation — zero patches to PVE system files.
Discovered automatically by Proxmox VE at runtime via Module::Load.

Implements:
- alloc_image:        create zvol via TrueNAS API + wire iSCSI extent/targetextent
- free_image:         tear down iSCSI association + delete zvol
- list_images:        enumerate zvols under configured pool/dataset
- status:             pool total/used/free via TrueNAS API (replaces SSH)
- path:               resolve /dev/disk/by-path from LUN ID via API
- activate_storage:   resolve target IQN + iscsiadm login
- deactivate_storage: iscsiadm logout + clear state cache
- activate_volume:    ensure session + wait for block device
- volume_has_feature: declare copy and snapshot support

iSCSI target resolution supports both explicit IQN (truenas_target) and
auto-discovery from TrueNAS portal config — the field being set or blank
acts as the gate between modes.

No SSH keys required. Bearer token auth only. REST API v2.0.
Transport: CORE 13.x + SCALE <= 24.10. WebSocket (SCALE 25.04+) in v3.1.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:31:39 -04:00
Kevin Adams 33faf6eb34 chore: v3.0 packaging cleanup — no patches, one .pm file
- Rename Custom/FreeNAS.pm → Custom/TrueNAS.pm (matches package name)
- Rewrite postinst: copy TrueNAS.pm + restart pveproxy (no patch logic)
- Rewrite postrm: remove TrueNAS.pm + restart (no restore-orig logic)
- Delete triggers: nothing to watch (no ZFSPlugin/pvemanagerlib/apidoc)
- Drop 'patch' dep, add 'open-iscsi'; update package description
- Delete patch-generation runbook (obsolete)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 19:55:59 -04:00
Kevin Adams fbfb3f26e1 chore: strip v2.x patch infrastructure — v3.0 clean slate
Remove everything that existed solely to wedge into PVE system files:
- stable-5 through stable-8 versioned patch archives
- pve-manager/js patch files (no more pvemanagerlib.js patching)
- pve-docs/api-viewer patch files (no more apidoc.js patching)
- perl5/PVE/Storage/LunCmd/ (replaced by PVE::Storage::Custom plugin)
- perl5/PVE/Storage/ZFSPlugin-*.patch (no more ZFSPlugin.pm patching)

Add perl5/PVE/Storage/Custom/FreeNAS.pm as the v3.0 starting point.
v3.0 ships one .pm file; PVE discovers it automatically — zero patches.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 19:50:55 -04:00