Commit Graph

225 Commits

Author SHA1 Message Date
Kevin Adams feb1776aa1 docs: update CLAUDE.md to reflect v3.0 single-repo build
Remove all references to the external packer repo and v2.x patch
approach. Document the self-contained build.yml pipeline, packaging
layout, and current open issues accurately.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 17:45:35 -04:00
Kevin Adams eb058266a9 docs: mark CORE 13.0-U6 and SCALE Electric Eel (24.10) as tested in v3.0 matrix
Cobia and Dragonfish remain listed as compatible but are not yet tested
against v3.0 — would require reinstalling those versions in the lab.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 17:10:11 -04:00
Kevin Adams e286929535 chore: remove v2.x files not needed in v3.0
REST/Client.pm — v3.0 uses LWP::UserAgent directly
stable-6/, stable-7/ — PVE 6/7 EOL, v3.0 requires PVE 8+
perl5/PVE/Storage/LunCmd/ — v2.x wedge approach replaced by Custom plugin
pve-manager/js/ patch files — v3.0 uses ui/truenas-storage.js instead

FreeNAS.pm and REST/Client.pm history preserved on master and release/2.x.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 17:02:05 -04:00
Kevin Adams 415c2a250b fix: unmap targetextent before extent delete in free_image (#259)
When a VM has multiple disks on the same per-VM target, migrating one
disk while the VM is running leaves the target in use for the other
disks. TrueNAS refuses to delete an extent associated with an in-use
target even with force=true.

Fix: delete the targetextent (LUN mapping) first, then delete the
extent. Removing the mapping severs this disk's association without
touching the active session or other LUNs on the same target.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:37:33 -04:00
Kevin Adams 5e2e5a8cbe fix: allow free_image on detached disks while VM is running (#258)
Removed _vm_is_running() check from free_image. When a disk is detached
from a running VM (unused0), QEMU immediately closes its libiscsi
connection — there is no active session to block the delete. The
force=true flag on the TrueNAS extent DELETE is the correct guard for
any residual session. The VM-running check was overly broad and blocked
legitimate deletes of already-detached disks.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:28:32 -04:00
Kevin Adams a4eb3e5557 docs: split Prerequisites by version — v3.0 needs no SSH keys or pre-created targets
v3.0 is fully API-driven: no SSH keys, no pre-created iSCSI targets.
v2.x still requires both. Separating them prevents confusion for users
on either version and removes incorrect guidance from the v3.0 path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:22:06 -04:00
Kevin Adams 678bf7a850 docs: add troubleshooting layer guide to distinguish plugin vs iSCSI vs Proxmox errors
Closes #257. Many support issues are misrouted because users don't know
which layer generated the error — the plugin (REST API orchestration),
the iSCSI/QEMU data path, or the Proxmox core storage stack. Added a
new section at the top of Troubleshooting that explains all three layers
with example log lines and a quick-triage table.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:15:28 -04:00
Kevin Adams d7b0735902 refactor: per-VM iSCSI targets with iscsi:// paths (#255)
Each VM now gets its own TrueNAS iSCSI target (proxmox-vm-<vmid>)
instead of sharing a single storage-wide target.  path() returns
iscsi://portal/iqn:proxmox-vm-<vmid>/lun — QEMU connects via libiscsi,
the same pattern ZFSPlugin.pm uses.

When a VM stops, QEMU closes its libiscsi connection.  TrueNAS sees
no active session on that VM's target, so extent DELETE with force=true
succeeds without stopping the iSCSI service.  Deleting one VM's disk
while other VMs are running on the same storage now works correctly.

Changes:
- Remove all iscsiadm session management (_iscsi_ensure_session,
  _iscsi_session_exists, _wait_for_device, _dev_path,
  _running_vms_on_storage, _resolve_target)
- Add _resolve_vm_target: find-or-create proxmox-vm-<vmid> target,
  inherit portal/initiator groups from existing targets
- Add _maybe_cleanup_vm_target: delete empty VM target after last disk
- Add _vm_is_running: check owning VM before deletion (uses PID file)
- Add _api_global: cached iSCSI global config (basename)
- path(): returns iscsi://portal/per-vm-iqn/lun (target looked up by
  target_id from extent, so legacy shared-target disks still work)
- activate_storage: API reachability check only
- activate_volume: verify extent exists, QEMU handles the connection
- deactivate_storage/volume: clear cache, return 1
- free_image: check owning VM stopped, force-DELETE extent, zvol,
  cleanup empty target — no service-stop fallback needed

Closes #255

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 15:04:01 -04:00
Kevin Adams 01318919fe fix: untaint PID in _running_vms_on_storage for PVE taint-mode Perl
PVE runs pvedaemon with -T (taint mode). Data read from files is tainted
and cannot be passed to kill() without explicit untainting. Extract PID
via regex capture which Perl treats as safe.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:35:46 -04:00
Kevin Adams 5684b335a1 fix: pre-flight check before service-stop fallback in free_image (#253)
Stopping the TrueNAS iSCSI service to clear session state (required on
CORE 13 when force=true is ignored) disrupts all LUNs on the target,
causing io-error on any other running VM using this storage.

Before triggering the service-stop path, scan /etc/pve/qemu-server/*.conf
for running VMs (PID file exists + process alive) that reference this
storage ID. If any are found, fail with a clear message naming the
blocking VMs rather than silently disrupting them.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:29:14 -04:00
Kevin Adams 178088197f fix: service-stop fallback for free_image on CORE 13 (#251)
force=true on extent/targetextent DELETE is unconditionally ignored by
CORE 13.0 — it always checks for active sessions. POST /service/restart
is async and iscsid reconnects in milliseconds, so the restart window
was never clear enough for the DELETE to succeed.

New fallback sequence (only triggered when force=true fails):
1. Set iscsiadm node to manual startup — prevents auto-reconnect
2. iscsiadm --logout — drops initiator session
3. POST /service/stop + poll until service is stopped (up to 15s)
4. DELETE targetextent — succeeds with service fully stopped
5. DELETE extent
6. POST /service/start — restore service
7. Restore automatic startup + discovery + login + rescan

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 11:20:11 -04:00
Kevin Adams c1f021c5a6 fix: hybrid free_image fallback — service restart + explicit targetextent delete (#251)
force=true on extent DELETE is not sufficient when CORE 13 has an active
or recovering session holding the targetextent (422 persists through retries).

Fallback path (triggered only when force fails and a targetextent exists):
1. POST /service/restart — purges all server-side session state
2. DELETE /iscsi/targetextent/id/{id} — now succeeds with service down
3. DELETE /iscsi/extent/id/{id} — clean delete with no association

Fast path (force=true) is tried first and handles the normal case without
any service disruption.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:24:36 -04:00
Kevin Adams 9cd3cfa234 refactor: remove initiator logout/login from free_image
Session management was only needed to prevent 422 when explicitly
deleting the targetextent.  Now that we delete the extent first with
force=true (which cascades targetextent removal server-side), the
initiator session can stay up throughout — same as the original
LunCmd/FreeNAS.pm delete path which never touched iscsiadm.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:09:55 -04:00
Kevin Adams 187f3f9e3f docs: explain GiB vs GB disk size discrepancy in Troubleshooting
Proxmox enters and allocates disk sizes in GiB (base-2) while TrueNAS
and many storage UIs display sizes in GB (base-10). This causes the
reported size to appear ~7% larger than the number entered, which is
correct and expected behaviour. Adds a reference table for common sizes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:06:02 -04:00
Kevin Adams 377503f89d fix: follow original FreeNAS.pm delete order — extent with force cascades targetextent
The original LunCmd/FreeNAS.pm (v2.0 path) deleted the extent with
{ force: true } and skipped the explicit targetextent DELETE entirely,
relying on TrueNAS to cascade it.  Our v3.0 had the order backwards:
deleting the targetextent first always hit 422 "target in use" when
any iSCSI session was active.

New approach mirrors the original:
1. Log out initiator session
2. DELETE extent with { force: true } — TrueNAS cascades targetextent
3. If that fails (CORE 13 strict session enforcement), restart iSCSI
   service and retry the extent DELETE (up to 5 times)
4. Restore session

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 09:03:38 -04:00
Kevin Adams 594b533a06 fix: volume_size_info override + free_image retry loop + remove rootdir content type
volume_size_info: base class routes through filesystem_path() → get_subdir()
which dies on block storage. Override to query TrueNAS API directly for
volsize.parsed so create_efidisk and other callers get the correct zvol size.

free_image retry loop: TrueNAS CORE 13 refuses targetextent deletion while any
iSCSI session is active. Single logout+restart was racy — pvedaemon workers
reconnect between the restart and DELETE when multiple disks are deleted
concurrently (e.g. VM destroy). Retry up to 5 times with fresh logout+service
restart each cycle and exponential backoff.

Remove rootdir from plugindata content types: rootdir signals LXC
container/directory storage and caused PVE to route TPM state allocation to
TrueNAS, which immediately failed since TrueNAS has no filesystem path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 08:57:25 -04:00
Kevin Adams f521089e57 fix: parse_volname + free_image session/LUN teardown for CORE 13
parse_volname: implement for vm-/base- volume names so PVE can resolve
device paths and hotplug disks. Without this the base class falls through
to directory-volume parsing and fails with a 400 hotplug error.

free_image: TrueNAS CORE 13 holds iSCSI sessions in recovery state after
TCP disconnect, blocking targetextent deletion with 422 even seconds after
logout. Fix: log out the initiator session by SID, then restart the TrueNAS
iSCSI service to immediately clear server-side session state. Delete the
targetextent and extent, then restore the initiator session so other LUNs
on the same target remain accessible.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 08:12:41 -04:00
Kevin Adams c6f662b88c fix: rescan iSCSI session in _wait_for_device to discover new LUNs
After TrueNAS reloads its iSCSI service, the PVE host's existing session
does not automatically pick up new LUNs. Adding iscsiadm --rescan every 5s
during the device wait loop lets the kernel discover newly exported LUNs
without requiring a full session logout/login.

Fixes the 30s hotplug timeout seen when adding a disk to a running VM.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 22:28:26 -04:00
Kevin Adams 835648fd69 fix: use /pool/dataset for space stats; default shared=1 for iSCSI
TrueNAS CORE 13.0 /pool API does not expose top-level size/free/allocated
fields (they are nested in topology.data[].stats). Switch status() to query
/pool/dataset?id=<pool> which has available.parsed + used.parsed on both
CORE and SCALE.

Add shared=1 as the default in the UI panel — iSCSI is a network block
device accessible from all cluster nodes, so it should be shared storage
by default.

Verified on pve01-hq against Tank01 (CORE 13.0-U6.7):
  TrueNAS01-Tank01  active  1804599296  165936464  1638662832  9.20%

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 22:17:18 -04:00
Kevin Adams 9c85fdb6e9 fix: extract pool name from path in status(); improve Pool/Dataset UX
status() was comparing the full dataset path (e.g. tank/proxmox/vdisks)
against TrueNAS /pool names which are top-level only (e.g. tank).
Extract the first path component so pool stats resolve correctly.

Rename Pool field label to 'Pool / Dataset Path' and update its hint
to make clear it accepts the full ZFS path (matching v2.x 'pool' field).
Rename Dataset to 'Sub-dataset' with a hint that discourages filling it
unless you genuinely need an extra sub-level beyond what's in Pool.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:52:41 -04:00
Kevin Adams 132875a5c5 fix: replace /iscsi/targetgroup with target.groups[] for CORE compatibility
TrueNAS CORE 13.0 does not expose /iscsi/targetgroup as a REST endpoint
(returns 404). Both CORE and SCALE include the portal group associations
inline in each target's 'groups' array from GET /iscsi/target, which is
all we need to filter targets by reachable portal IP.

Remove the /iscsi/targetgroup call entirely — use target.groups[].portal
cross-referenced against /iscsi/portal results instead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:42:19 -04:00
Kevin Adams c5868dd4f6 fix: remove sensitive-properties — store API key in storage.cfg
PVE's sensitive-properties mechanism extracts listed keys from $param
before check_config and passes them only to on_add_hook/on_update_hook.
activate_storage reads from $cfg which never receives those values, so
the API key was always missing at runtime.

The API key now lives in storage.cfg (root-readable, mode 0640, same as
the v2.x truenas_secret field). Proper on_add_hook private-file storage
is tracked in issue #247.

Restore truenas_api_key => {} (required on create). PVE's update flow
passes $create=0 to check_config which skips absent keys, so
edit-without-changing-key still works.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:34:36 -04:00
Kevin Adams efcad405e5 fix: sensitive-property handling + UI autofill + API key reveal button
truenas_api_key must be optional in options() because PVE extracts
sensitive-properties from the POST body before calling check_config,
causing a spurious 'missing required option' 500 on storage create.
Add a guard in _api() so a missing key produces a clear error.

UI: add autocomplete="url" on host and "new-password" on API key to
prevent Firefox/Chrome from filling in PVE login credentials.
Add a reveal trigger button to show/hide the API key field.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 21:17:58 -04:00
Kevin Adams ea1c2208e1 chore: update perlcriticrc — target TrueNAS.pm, clean up stale comments
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:58:14 -04:00
Kevin Adams 8293a718a3 ci: rewrite build pipeline for v3.0 — no patches, TrueNAS.pm + JS only
Remove the validate-patches job (no more patch files in v3.0).
Update lint to target TrueNAS.pm instead of FreeNAS.pm.
Rebuild staging assembly: TrueNAS.pm + truenas-storage.js are the
only payload — no patch dirs, no REST-Client.pm, no triggers file.
Add $VERSION = '3.0.0' to TrueNAS.pm as the single source of truth.
Add release/3.x as a testing-channel branch alongside master.
Pin actions to @v4 (upload/download-artifact, checkout).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:57:18 -04:00
Kevin Adams 8a487ef4a8 feat: add TrueNAS UI panel via injected JS — no pvemanagerlib.js patch
Instead of diffing and patching the 50k-line pvemanagerlib.js each PVE
release, we ship truenas-storage.js as our own file and inject one
<script> line into index.html.tpl (rarely changes, ~60 lines).

truenas-storage.js adds 'truenas' to PVE.Utils.storageSchema at runtime
and defines PVE.storage.TrueNASInputPanel with all config fields:
  - TrueNAS host, API key, pool, dataset (column 1)
  - SSL toggle, cert verify, portal IP, target IQN (column 2)
API key field: required on create, optional on edit (only submitted if
the user types a new value).

postinst: copies TrueNAS.pm + truenas-storage.js, adds script tag
postrm:   removes both files, strips script tag

Closes #225 (core interface), #226 (iSCSI lifecycle activate/deactivate),
#227 (API pool listing replaces SSH).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:50:15 -04:00
Kevin Adams f6ca0c9ef5 fix: correct package name and API version for PVE 8.4.x
Package must match filename (TrueNAS.pm → PVE::Storage::Custom::TrueNAS)
so PVE's module loader finds the class when checking ISA PVE::Storage::Plugin.

api() returns 11 to match APIVER in PVE::Storage on PVE 8.4.x.
Returning 10 loaded successfully but triggered a deprecation warning.

Verified: pvedaemon loads plugin cleanly with no errors or warnings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:36:13 -04:00
Kevin Adams 432be1ba5c feat: implement TrueNAS v3.0 custom storage plugin (#225, #226, #227)
Full PVE::Storage::Custom implementation — zero patches to PVE system files.
Discovered automatically by Proxmox VE at runtime via Module::Load.

Implements:
- alloc_image:        create zvol via TrueNAS API + wire iSCSI extent/targetextent
- free_image:         tear down iSCSI association + delete zvol
- list_images:        enumerate zvols under configured pool/dataset
- status:             pool total/used/free via TrueNAS API (replaces SSH)
- path:               resolve /dev/disk/by-path from LUN ID via API
- activate_storage:   resolve target IQN + iscsiadm login
- deactivate_storage: iscsiadm logout + clear state cache
- activate_volume:    ensure session + wait for block device
- volume_has_feature: declare copy and snapshot support

iSCSI target resolution supports both explicit IQN (truenas_target) and
auto-discovery from TrueNAS portal config — the field being set or blank
acts as the gate between modes.

No SSH keys required. Bearer token auth only. REST API v2.0.
Transport: CORE 13.x + SCALE <= 24.10. WebSocket (SCALE 25.04+) in v3.1.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 20:31:39 -04:00
Kevin Adams 33faf6eb34 chore: v3.0 packaging cleanup — no patches, one .pm file
- Rename Custom/FreeNAS.pm → Custom/TrueNAS.pm (matches package name)
- Rewrite postinst: copy TrueNAS.pm + restart pveproxy (no patch logic)
- Rewrite postrm: remove TrueNAS.pm + restart (no restore-orig logic)
- Delete triggers: nothing to watch (no ZFSPlugin/pvemanagerlib/apidoc)
- Drop 'patch' dep, add 'open-iscsi'; update package description
- Delete patch-generation runbook (obsolete)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 19:55:59 -04:00
Kevin Adams fbfb3f26e1 chore: strip v2.x patch infrastructure — v3.0 clean slate
Remove everything that existed solely to wedge into PVE system files:
- stable-5 through stable-8 versioned patch archives
- pve-manager/js patch files (no more pvemanagerlib.js patching)
- pve-docs/api-viewer patch files (no more apidoc.js patching)
- perl5/PVE/Storage/LunCmd/ (replaced by PVE::Storage::Custom plugin)
- perl5/PVE/Storage/ZFSPlugin-*.patch (no more ZFSPlugin.pm patching)

Add perl5/PVE/Storage/Custom/FreeNAS.pm as the v3.0 starting point.
v3.0 ships one .pm file; PVE discovers it automatically — zero patches.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 19:50:55 -04:00
Kevin Adams 6ec31f76a7 docs: fix TrueNAS CORE 13 API key navigation path
CORE 13 moved API Keys under the gear icon (top-right), not System menu.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 22:53:25 -04:00
Kevin Adams e9907d0041 feat: deprecate basic auth with syslog warning (#244)
Basic auth (username/password) has been supported since the plugin's
origin but TrueNAS API keys have been available since TrueNAS 12 (2020).
Sending credentials on every API call is a security risk and TrueNAS
may stop accepting it in future releases.

Log a syslog(warn) on every connection attempt that uses Basic Auth so
operators see the deprecation message in their Proxmox logs. The warning
links to issue #244 which tracks removal in v3.0.0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 22:42:39 -04:00
Kevin Adams ee708f3148 docs: add SCALE 25.04 API key revocation warning and update volblocksize note
- New troubleshooting entry: SCALE 25.04 revokes existing API keys on
  upgrade; users must regenerate and update truenas_secret in PVE config.
  Also notes HTTPS enforcement and links to #243 for REST deprecation.
- Updated volblocksize entry: v2.4.0 auto-corrects blocksize, so the
  "fix" is now informational rather than a manual step.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 22:35:06 -04:00
Kevin Scott Adams c166501873
feat: auto-detect and correct ZFS blocksize for TrueNAS SCALE/CORE (#241)
* feat: auto-detect and correct ZFS blocksize for TrueNAS SCALE/CORE (#241)

SCALE requires >= 16k volblocksize; CORE works with 8k. Without correction,
creating a disk on SCALE with blocksize=8k triggers a ZFS warning and wastes
space on SCALE's minimum-4k-block pool layout.

Three changes:
- freenas_get_recommended_blocksize: fixes freenas_api_connect → freenas_api_check
  so product_name is actually populated before the SCALE check (was always returning
  8192 before this fix)
- freenas_parse_blocksize: new helper to compare "8k"/"16k"/integer blocksize strings
- alloc_image (both 8.x and 8.4.x patches): always detect recommended blocksize for
  freenas provider; if configured < recommended, override for this call AND persist
  the correction back to storage.cfg via lock_storage_config
- on_add_hook (8.4.x patch only): detect at storage creation time and correct $scfg
  before write_config saves it — no extra write needed

Tested on pve01-hq (PVE 8.4.19) against Tank02 (SCALE 24.10.2.1):
- Disk created with blocksize=8k → task log shows correction message, storage.cfg
  updated to 16384, no volblocksize warning from TrueNAS
- Disk created with blocksize=16384 → no-op, clean TASK OK

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* ci: add validate-patches step for ZFSPlugin PVE 8.4.x patch

ZFSPlugin-8.4.14_1.pm.patch has been in the build since v2.3.0 but was
never dry-run validated in CI. Now that we have the 8.4 orig committed
(ZFSPlugin-8.4.14_1.pm.orig from libpve-storage-perl 8.3.8), wire it up.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: replace explicit return undef with bare return (perlcritic)

Subroutines::ProhibitExplicitReturnUndef violation in
freenas_get_recommended_blocksize. Bare return in list context returns
an empty list rather than a list containing undef, which is the
correct Perl idiom.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 11:49:05 -04:00
Kevin Adams a0d0285194 ci: fix versioning — stable uses -1 revision, pre-releases use ~ prefix
Stable releases now produce X.Y.Z-1 (e.g. 2.3.0-1) so they sort higher
than any pre-release build in dpkg. Pre-release builds now use tilde
(~beta+sha, ~alpha+sha) which sorts below the base version, ensuring
apt upgrade always selects stable over a previously installed beta.

Previously -beta+sha sorted higher than X.Y.Z causing apt to refuse
upgrades from beta to stable.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 16:25:34 -04:00
Kevin Adams aa617cf788 ci: upgrade softprops/action-gh-release to v3 and pin cloudsmith action
softprops/action-gh-release@v3 adds Node 24 support (v2 was Node 20).
Pin cloudsmith-io/action to v0.6.14 (latest stable) instead of @master.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 15:59:19 -04:00
Kevin Scott Adams 823ef14139
feat: v2.3.0 — hardening, stability, and PVE 8.4.x support
v2.3.0: PVE 8.4.x compatibility, packaging consolidation, clone/migrate fixes
2026-05-20 12:25:56 -04:00
Kevin Adams 6571ab23d1 fix: remove zvol delete from run_delete_lu to eliminate false negative (#240)
ZFSPlugin's free_image deletes the zvol via SSH after calling our
run_delete_lu. Our earlier addition of freenas_delete_zvol in
run_delete_lu caused a double-delete: SSH would fail, free_image's
error recovery would call run_create_lu on the now-missing zvol,
producing a spurious "Unable to create lun (rolled back)" in the task
log even though the migration had already succeeded.

freenas_delete_zvol remains in run_create_lu's rollback path where it
correctly cleans up zvols orphaned by a failed LUN creation (#239).

Also documents ZFS blocksize requirements for TrueNAS SCALE (16k) vs
CORE (8k) in README configuration and troubleshooting sections (#241).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 11:50:12 -04:00
Kevin Adams 267e14ebfa fix: explicitly delete zvol via pool/dataset API on extent removal (#239)
remove:true in the iSCSI extent DELETE body is unreliable on TrueNAS
SCALE 24.10+ — the extent config is removed but the underlying zvol is
left behind as an orphaned dataset. Add freenas_delete_zvol() which calls
DELETE /api/v2.0/pool/dataset/id/{path}/ explicitly after every extent
removal, covering both the rollback path in run_create_lu and normal
disk deletion in run_delete_lu. Guarded to v2.0 API only using the
per-host cache to avoid stale package-global state when two storage
hosts with different API versions are active simultaneously.

Verified on TrueNAS SCALE 24.10.2.1: zvol is fully removed from
Datasets after a forced rollback test (FREENAS_TEST_ROLLBACK=1).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 08:56:02 -04:00
Kevin Adams ec4085b244 fix: full eval-based rollback on LUN creation failure, add test trigger (#239)
Replace the ad-hoc if/rollback in run_create_lu with an eval block that
tracks both extent_id and link_id so either resource is cleaned up if
anything fails at any point during creation (#214 only covered the
link-creation step).

Also adds a FREENAS_TEST_ROLLBACK env-var trigger: set it in pvedaemon's
environment to force a rollback after a successful create, allowing
verification that orphaned extents and links are removed without needing
a real failure condition.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 00:00:55 -04:00
Kevin Adams 168cf0adf7 fix: repair freenas_api_connect autovivification and ping URL passthrough (#238)
Two bugs caused clone/migrate to fail with 'setHost on undefined value'
followed by 'Loop recursion prevention':

1. freenas_api_check reads $freenas_server_list->{$apihost}{client} to
   check if initialized. Perl autovivifies {$apihost} as an empty {} on
   that read. freenas_api_connect then saw a defined (but empty) hash and
   skipped initialization, leaving client undef → setHost crash.
   Fix: guard also checks !defined ...{client}.

2. $ping was a local variable reset to v1.0 on each recursive call.
   TrueNAS 25.04+ removed the v1.0 endpoint (returns 404), so the
   v1.0→v2.0 upgrade branch looped until the runaway counter tripped.
   Fix: accept $ping as optional second parameter (defaults to v1.0)
   and pass it through on both recursive call sites.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 23:34:58 -04:00
Kevin Adams 3b3cc8d135 fix: ZFSPlugin 8.4.x patch and missing patch dependency (#236 #237)
Two install-time failures on PVE 8.4.x nodes:

1. 'patch' utility not in Depends — causes silent patch failure on fresh
   nodes; all three PVE file patches silently skip, plugin appears to
   install but storage operations fail. Added 'patch' to control.j2.

2. ZFSPlugin.pm routing not applied on PVE 8.4.x — tabs→spaces reformat
   and get_base($scfg) signature change in 8.4 broke hunks #3 and #4
   (the freenas elsif routing branches). Nodes with a prior partial
   install are doubly affected: grep idempotency check passes on partial
   freenas strings, so the routing is never applied, causing:
     "freenas: unknown iscsi provider" on clone/migrate operations.

New ZFSPlugin-8.4.14_1.pm.patch generated against stock PVE 8.4.19,
bundled as patches/ZFSPlugin/8.4.patch alongside existing 8.patch.

Verified clean on pve01-hq (8.4.19) and pve02-hq (8.4.14).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 23:16:17 -04:00
Kevin Adams b936e42460 fix: correct pvemanagerlib 8.4 patch anchor to prevent wrong-panel injection
The initComponent anchor text was not unique in 8.4.19 — it matched
pveLxcCPUInputPanel before ZFSInputPanel, causing tnsecret/tnconfirmsecret
to be declared in the wrong scope. The ZFSInputPanel dialog then threw a
ReferenceError on open, preventing the modal from appearing.

Fixed by anchoring hunk 6 on the already-modified setValues block
(which is unique to ZFSInputPanel after hunk 5 is applied), ensuring
tnsecret/tnconfirmsecret always land in the correct initComponent scope.

Verified on 8.4.14 and 8.4.19 stock files. Tested live on pve01-hq (8.4.19).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 21:37:05 -04:00
Kevin Adams b41fefb72b docs: add PVE version compatibility warnings across all user touchpoints
Users on unsupported PVE versions now get clear guidance at every
layer — README, install time, and GitHub release notes — rather than
silently failing or getting confusing patch errors.

- README: replace bare compatibility table with a full matrix and
  prominent callout blocks for PVE 7 (last v2.x), PVE 8 (EOL
  2026-08-31), PVE 9+ (use v3.0), and PVE ≤6 (unsupported)
- postinst: add check_pve_version() that exits on PVE < 7, warns
  loudly on PVE 7 (stay on v2.x, do not upgrade to v3.0), notes
  the PVE 8 EOL date, and warns on PVE 9+ (v2.x not supported)
- build.yml release template: add a version compatibility table so
  every GitHub release prominently states who should and should not
  install it

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 18:33:13 -04:00
Kevin Adams ea1d0e2248 fix: version-aware pvemanagerlib patch selection for PVE 8.4.x (#223)
PVE 8.4 reformatted pvemanagerlib.js (tabs→spaces, structural changes),
making the existing 8.0.5-era patch fail on all 8.4.x nodes.

- Add pvemanagerlib-8.4.14_1.js.patch covering PVE 8.4.x (all hunks
  verified on 8.4.14 and 8.4.19)
- Update postinst to detect pve-manager major.minor (e.g. "8.4") and
  try a minor-version-specific patch before falling back to the major
  patch — 8.3.x nodes keep using 8.patch, 8.4.x nodes use 8.4.patch
- Update build.yml to bundle 8.4.patch alongside the existing 8.patch

Tested live on pve01-hq (8.4.19) and pve03-hq (8.3.2) — correct patch
selected in each case, pvemanagerlib patched successfully on 8.4.19.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 18:10:11 -04:00
Kevin Adams 332b6a8257 docs: add PVE 8.4/9.0/9.1 research brief and ADR-008 version support matrix
Reviewed PVE 8.4, 9.0, and 9.1 release notes for storage subsystem changes
relevant to the plugin and Phase 3 design.

Key findings:
- iSCSI hostname portal support (PVE 8.4+) — gap in current plugin (#233)
- Snapshot-as-Volume-Chains (PVE 9.0 tech preview) — Phase 3 opportunity (#234)
- GlusterFS dropped in PVE 9 — confirms need for proper custom plugin
- pvemanagerlib.js still broken on PVE 8.4.x (#223, blocked)

ADR-008 records the PVE version support matrix:
- v2.x: PVE 7 (best-effort), PVE 8 (supported), PVE 9 (not supported)
- v3.0: PVE 8+ (core), PVE 9.0+ (snapshots)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 13:05:08 -04:00
Kevin Adams 4833cfbcca docs: finalize ADR statuses to Accepted; add ADR README and KSA scaffold
All seven ADRs were in draft/proposed/decided states — promote all to
Accepted now that the decisions are implemented and stable.

Add .claude/cos/adrs/README.md documenting the ADR lifecycle process.

Add KSA Workspace Standards section to CLAUDE.md (sourced from scaffold
at claude-scaffold/CLAUDE.md) so future sessions inherit workspace-wide
rules without needing a separate file load.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 11:07:49 -04:00
Kevin Scott Adams 9876e85d74
Merge pull request #232 from TheGrandWazoo/ci/upgrade-actions-node24
ci: upgrade GitHub Actions to Node 24 runtimes (closes #231)
2026-05-19 11:03:27 -04:00
Kevin Adams 89acaf4145 ci: upgrade GitHub Actions to Node 24 runtimes (closes #231)
GitHub is removing Node.js 20 from Actions runners on Sep 16, 2026
and forcing Node 24 as default from Jun 2, 2026. Upgrade all
first-party actions to latest major versions that ship with node24:

  actions/checkout@v4        → v6
  actions/upload-artifact@v4 → v7
  actions/download-artifact@v4 → v8

Note: setup-python@v5 warning is inside cloudsmith-io/action@master
and is not directly addressable here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 10:06:20 -04:00
Kevin Scott Adams 4edd50b300
Fix dangling extent on failed LUN creation (issue #214)
Fix dangling extent on failed LUN creation (issue #214)
2026-05-18 17:14:12 -04:00