The post-ADR-014 multipath re-check was only run as a standalone script,
the exact testing gap that caused #290 in the first place. Closed it by
re-verifying through a real pveproxy/pvedaemon-driven VM disk allocate +
start + teardown cycle against .92, and documented the result in ADR-014
and architecture.md §10 alongside the original AnyEvent-era verification.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AnyEvent::WebSocket::Client's blocking ->recv() cannot run inside
pveproxy/pvedaemon (themselves built on AnyEvent as their core
reactor) without nesting event loops, which AnyEvent refuses by
design -- confirmed by a real user on the first real WebUI use of
v4.0.0, one day after release. Not a race condition: this fails on
every real "Add Storage" attempt against a WebSocket-transport host,
unconditionally. v4.0.0's testing never caught it because every
verification script ran standalone, never inside the actual daemon
process -- pvesm (a fresh CLI process) doesn't trigger it either, only
the real REST API path through pveproxy does.
Replaced _ws_connect/_ws_call's internals with Protocol::WebSocket::Client
+ IO::Socket::SSL -- a plain synchronous socket client with zero
event-loop dependency, so there's no shared reactor state to conflict
with pveproxy's. _api_ws's method-mapping logic is completely
unchanged; only the connection/call internals were rewritten.
Reproduced the exact bug via a real POST to /api2/json/storage through
pveproxy on a lab node, then confirmed the fix resolves it the same
way. Re-ran every check from v4.0.0's release (read path, write path,
snapshots, multipath) against the new transport -- zero regressions.
Full story: ADR-014, superseding ADR-012's library-choice section
only (everything else in ADR-012 stands). $VERSION bumped to 4.0.1.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Written after the v4.0.0 release housekeeping (CHANGELOG, ROADMAP,
issue closing) happened as an afterthought rather than before the
tag was pushed. Discipline checklist, not automatable -- the ROADMAP
write-up and issue-closing narrative need synthesis a script can't do.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CHANGELOG's [Unreleased] section becomes a dated [4.0.0] entry;
ROADMAP's "Upcoming" section moves to "Released", matching the
pattern every prior version got.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Full volume_snapshot/_info/_rollback/_delete cycle against .92
succeeded with zero param-shape issues on the first try. Every API
call this plugin makes over the WebSocket transport is now
live-verified. See ADR-012.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Everything else in the write path (target/extent/targetextent/dataset
create+delete) is now live-confirmed, see ADR-012.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Live-tested the write path against .92 for the first time (alloc_image
-> path -> volume_size_info -> free_image full cycle) and found two
real bugs in the inferred delete mappings:
- iscsi.target.delete and iscsi.targetextent.delete were passing a raw
regex-captured id as a Perl string; Pydantic rejects "Input should
be a valid integer" without int().
- iscsi.extent.delete(id, remove=False, force=False) is positional --
REST's DELETE body {force=>true} was being passed as a single hash
in position 2, which the API validates as a boolean field named
"remove", not force. Fixed to unpack positionally: id, remove
(always false -- REST never asked for this), force (from the REST
body's force field).
Confirmed clean after the fix: a full alloc_image/free_image cycle
against .92 leaves zero orphaned targets, extents, targetextents, or
datasets. This is the first write-path confirmation for the WebSocket
transport -- narrows ADR-012's "not yet live-tested" warning to the
remaining untested calls (targetextent.create/query filters beyond
what alloc_image exercises, pool.dataset.delete's options shape,
zfs.snapshot.* entirely).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds _api_ws/_ws_connect/_ws_call and a _transport() auto-detector so
_api() dispatches to REST or WebSocket per host based on the detected
TrueNAS version -- storage.cfg is unchanged either way. Follows
ADR-012's decided design: AnyEvent::WebSocket::Client for the
transport, auth.login_with_api_key (not auth.login_ex, which needs a
username this plugin has no way to know), and the JSON-RPC envelope
hand-built with the existing JSON module rather than a dedicated RPC
library.
Live-verified end-to-end against .92 (TrueNAS-25.04.2.6, real SCALE
box) through the actual plugin module, not just the diagnostic script:
transport auto-detection, _api_global, portal/target queries, and the
single-ID lookup mapping used by path()/qemu_blockdev_options() all
confirmed working. This exercises every *read* (query) call the
plugin makes.
Caught and fixed one real bug during integration testing:
/system/product_type returns the license tier (e.g.
"COMMUNITY_EDITION"), not the product family -- CORE vs SCALE
detection uses the /system/version string's major-version magnitude
instead (SCALE's calendar versioning vs CORE's legacy 11.x-13.x range).
Scope note (see file header): the *write* path (create/delete/update
param shapes) follows TrueNAS's documented CRUDService conventions but
has not yet been independently live-tested -- verify against .91/.92
before trusting alloc_image/free_image/snapshot operations in
production on SCALE 25.04+.
$VERSION bumped to 4.0.0. Adds libanyevent-perl and
libanyevent-websocket-client-perl as package + CI lint dependencies.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Read-only tool for verifying ADR-012's transport assumptions against a
real TrueNAS SCALE 25.04+ host before any of this lands in TrueNAS.pm.
Not part of the packaged plugin. Used to confirm auth.login_with_api_key
and the full method-mapping table live against a Goldeneye (25.10.3.1)
lab node.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- CHANGELOG.md: new [Unreleased] section tracking v4.0.0 progress (added/
verified/fixed/deferred), per this project's rule to keep docs current
as work lands rather than batching into eventual release notes
- docs/architecture.md: new §5 "API Flow -- TrueNAS WebSocket JSON-RPC 2.0
(v4.0.0+)", inserted right after the REST API Flow section for narrative
continuity -- covers transport auto-detection, the /system/product_type
bug, connection/auth, the JSON-RPC envelope and JSON::PP::Boolean gotcha,
the full method-mapping table with both confirmed-broken positional-arg
cases called out, and multipath's free inheritance. Sections 5-10 renumbered
to 6-11 accordingly (TOC, cross-references, and the old release/N.x branch
workflow references all updated)
- Also fixed unrelated staleness found along the way: "Branch from master"
in the Contributing section (pre-ADR-013), and the ADR quick-reference
table, which was missing ADR-009 through ADR-013 entirely and still
described ADR-008 as current despite it being superseded by ADR-009
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Full alloc_image/activate_volume/path/deactivate_volume/free_image
cycle against .92 with zero code changes to TrueNASMultipath.pm --
real dm-multipath device came up with 2 active paths. Also notes a
practical finding: .92 has only one TrueNAS portal object (0.0.0.0)
rather than two per-IP ones, and iscsiadm still produced two genuine
independent paths logging into different IPs against it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ran the full read/write/snapshot suite against .90 (TrueNAS CORE,
real production box) to confirm _transport()'s dispatcher wrapper
doesn't regress the REST path for hosts that never touch WebSocket.
Zero regressions, confirmed back to original state afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every read and write API call the WebSocket transport makes is
live-verified against real TrueNAS SCALE hosts. Kevin's sign-off.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every API call the WebSocket transport makes is now live-verified
against a real TrueNAS SCALE host with zero param-shape issues
remaining. Left as Draft pending Kevin's own sign-off, per his
explicit instruction earlier.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
alloc_image/path/volume_size_info/free_image confirmed working
end-to-end against .92 after fixing the delete param-shape bugs.
pool.dataset.create/delete both confirmed along the way. Only
zfs.snapshot.* remains entirely untested against the WS transport.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_api_ws/_ws_connect/_ws_call/_transport landed on next, live-verified
against .92 (25.04.2.6) through the real plugin module. Also notes
the /system/product_type bug caught during integration (returns
license tier, not product family -- fixed by parsing /system/version
magnitude instead).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Kevin rebuilt .92 from scratch specifically to TrueNAS-25.04.2.6 -- the
exact SCALE minor version NAS-135643 was originally filed against, two
patch levels later. iscsi.target.query returned a real target cleanly,
no error. This was the one item ADR-012 flagged as an acknowledged gap
rather than a completed check; no remaining live-verification gaps for
this ADR now. Also documents a minor auth-flow gotcha found along the
way: retrying auth.login_ex on the same connection after a failed
attempt threw an unhandled RuntimeError on this fresh install,
reinforcing the login_with_api_key recommendation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
release/3.x -> main, release/4.x -> next, master -> archive/v2-legacy
(retired, was byte-identical to release/2.x). Branch identity is now
role-based, not version-numbered -- main and next are fixed names
that never get renamed at a major-version cutover; release/N.x is
reserved exclusively for already-superseded majors, created only at
the moment they're archived.
This directly fixes the structural cause of the bug found 2026-08-30:
build.yml hardcoded literal branch-name comparisons for the
testing/beta channel, and ADR-006's original "master -> beta channel"
rule was never superseded when release/3.x took over as the active
branch, leaving master as a silent, stale duplicate of release/2.x
that CI still treated as a legitimate publish source. Rewrote the
version-resolution logic to match roles instead: refs/heads/main,
refs/heads/next, and a refs/heads/release/*.x pattern that auto-
catches any future archived major with zero code changes needed.
Also tightened the push/PR triggers from a bare wildcard to an
explicit allowlist, and switched the release-notes CHANGELOG link to
a tag-relative permalink instead of a floating branch reference.
Research (DEP-14, openedx/frontend-base #273, semantic-release's
branch-pattern config) is cited in ADR-013; the openedx precedent in
particular fixed this identical bug the same way.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.91 was upgraded straight to TrueNAS-25.10.6 (skipping the planned
Fangtooth stop) and re-tested clean: all 8 method-mapping families
work against real production-scale iSCSI config (2 targets, 9
extents), and NAS-135643 still doesn't reproduce. Also corrects the
auth.login_ex conclusion from the .92-only pass -- it works fine given
the right username (.91's key IS root, .92's isn't), so the
login_with_api_key recommendation stands for a different reason: no
username field exists in storage.cfg to use it correctly, not because
login_ex is broken. Notes the resulting gap (no 25.04.x node left in
the lab) and the revised lab-node role split.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Not committed to any release yet -- filed as a backlog idea alongside
today's #243/ADR-012 work, since a dual REST/WebSocket transport
reality in v4.0.0 makes visible version mismatches more useful to
surface.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two things surfaced since this ADR was last updated: confirmed
TrueNASMultipath.pm needs no separate WebSocket work (it calls the
shared _api() helper directly), and release/4.x now exists as the
implementation branch, including the dormant CI bug it exposed and
the fix for it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Multipath (truenas-proxmox-multipath, shipped v3.2.0/#256) had zero
coverage in README.md and was actively described as unsupported/future
work in docs/architecture.md and docs/migration-advanced.md. Added:
- README: Prerequisites, Installation, and Configuration sections for
the multipath package and its truenas-multipath storage type
- architecture.md: new §9 covering what TrueNASMultipath.pm overrides
vs. inherits, and the live failover behavior confirmed 2026-08-29
- migration-advanced.md: points to the multipath package instead of
calling it unsupported
- truenas-storage-help.html: single-path note pointing to the
multipath alternative
Also fixed README content that was stale from the v3.0 beta era: the
PVE 9 blockdev issue (#266) was fixed in v3.1.0 but still read as
open, "v3.0 (beta)" headings never updated to v3.x, a broken internal
anchor link, and a note claiming WebSocket support was "v3.0.0" scope
when ADR-009 placed it at v4.0.0.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The "Publish to GitHub Pages testing dist" step ran for both 'testing'
and 'development' channel builds and does rm -rf pool/testing before
copying in the new .deb — so any push to a release/* branch other than
release/3.x silently overwrote the real testing dist with its own alpha
build. Confirmed happening 2026-08-30 when release/4.x was created and
its 3.2.5~alpha+d0ea761 build replaced release/3.x's legitimate beta
build there. Restricted the step to channel == 'testing' only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ran a diagnostic against the .92 TrueNAS SCALE lab node (25.10.3.1) to
check ADR-012's open assumptions. Confirms auth.login_with_api_key over
auth.login_ex (the latter needs a username field this plugin doesn't
collect), verifies all 8 REST->JSON-RPC method mappings live, and finds
NAS-135643 no longer reproduces on this version. Also documents a real
Perl gotcha: JSON's decode_json returns a blessed JSON::PP::Boolean for
true/false, not a plain 1/0.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scopes the transport-abstraction design called for in #243: confirms the
JSON-RPC 2.0 method mapping for all 8 REST endpoint families the plugin
uses, decides on AnyEvent::WebSocket::Client (apt-packaged, no CPAN-only
deps) after finding an independent prior implementation, and defers the
#249 per-variant file split to a future v4.1/v5.0. Still Draft pending
live verification against the lab's Fangtooth/Goldeneye TrueNAS nodes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Four of the six tracked issues (#228, #234, #256, #260) had shipped
since this list was last touched, and #250 was closed too — the
instructions Claude reads at session start were actively wrong.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Debian trixie recommends deb822-style .sources files over the one-liner
.list format, and PVE 9 ships on trixie. Show deb822 as the primary form
in README (stable, v2 track, Cloudsmith migration, testing) and
getting-started.md/upgrade-paths.md, keeping the one-liner .list format
as a fallback for PVE 8 / bookworm users.
Fixes#285
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same root cause as #179: postinst/postrm only restarted pvedaemon/pveproxy,
never pvestatd/pve-ha-lrm/pve-ha-crm, so HA-managed VMs on truenas-multipath
storage could fail to start after install/upgrade. Mirrors the #179 fix —
bare statements so set -e aborts on a real failure, HA daemons restarted
via systemctl (they don't support a restart subcommand). Extended
test-restart-services.sh to cover both multipath scripts with the same
stub approach.
Fixes#287
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CHANGELOG.md stopped at 3.0.0 despite five releases since. Backfilled
from git/tag history so the changelog is a reliable release record again.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
README and getting-started.md still described the legacy v2.x "ZFS over
iSCSI" storage type and pre-25.04 TrueNAS SCALE UI locations. Update both
to match the actual v3 "TrueNAS (ZFS/iSCSI)" panel fields, document the
SCALE 25.04+ API key location and the throwaway-share workaround for the
now-missing standalone portal/initiator-group screen, and sync the help
HTML per the same-commit doc-sync rule.
Fixes#286
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
core.fileMode=false on this checkout (NTFS-mounted /mnt/c) meant the
previous commit recorded tests/stubs/* and tests/fail-stubs/* as 100644
despite being executable locally, so CI failed with "Permission denied"
trying to exec them via PATH lookup.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
pvedaemon/pveproxy picked up the patched PVE::Storage state, but
pve-ha-lrm/pve-ha-crm kept running against the pre-install state, so
HA-managed VMs on TrueNAS storage failed to start. They don't support
a `restart` subcommand (only start/stop/status), so they're restarted
via systemctl instead. Also switched every restart from `cmd && log
...` to bare statements — set -e does not fire on the left side of
&&, so a genuine restart failure was previously swallowed silently
instead of aborting the script.
Adds tests/test-restart-services.sh: sources the real postinst/postrm
against stub PVE binaries in CI to catch both classes of regression
(wrong subcommand, silently-swallowed failure) without needing a live
PVE cluster. Real HA-cluster validation is documented as a manual
runbook (RUNBOOK-001) since GitHub Actions can't run pve-ha-lrm/crm.
Bumps $VERSION to 3.2.4.
Fixes#179
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
pvedaemon/pveproxy picked up the patched PVE::Storage state, but
pve-ha-lrm/pve-ha-crm kept running against the pre-install state,
so HA-managed VMs on TrueNAS storage failed to start. pve-ha-lrm
and pve-ha-crm don't support a `restart` subcommand (only
start/stop/status), so they're restarted via systemctl instead.
Confirmed root cause by mir07/deamonkai on #179; verified the full
restart_pve_services() sequence end-to-end on pve01-hq before commit.
Fixes#179
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
SCALE 25.04 tightened the iSCSI portal PUT schema — sending port in the
listen array now returns "Extra inputs are not permitted". Added §6.9 in
getting-started.md with the correct curl example, and a matching row in
the help panel troubleshooting table.
Closes#280
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace bare pveversion with pveversion -v (includes full package list
for better triage). Replace grep on /var/log/syslog with journalctl
(syslog doesn't exist on Debian 12+/PVE 9).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Users who had the v3 repo configured before v3.2.2 see an apt security
prompt when the Codename field changed from unset to 'error'. Document
the --allow-releaseinfo-change one-liner in README, getting-started.md
(new section 6.8), and upgrade-paths.md (codename column + callout).
Closes#284.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
PVE 9.2.3 sets APIVER=15 in PVE::Storage. Our plugin was returning 14,
triggering "implementing an older storage API" on every pvedaemon/pveproxy
restart. No functional change — plugin loaded and worked correctly at 14.
Closes#282
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
dists/error was never written to gh-pages because the codename alias
code was added after the v3.2.0 tag. The stable publish job only runs
on v*.*.* tags, so all subsequent pushes to release/3.x only touched
dists/testing. Bumping to v3.2.1 triggers stable publish to create it.
Also fix: cp -r src dst when dst already exists puts src inside dst
rather than replacing it. Add rm -rf before cp -r to guarantee a clean
replace on every stable release.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Codename was added to Release metadata but dists/<codename>/ directories were
never created on GitHub Pages, so apt would 404 on InRelease when using the
codename as the dist name. Copy each dist dir under its codename alias after
the main loop so both Suite and Codename values work interchangeably.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
v2/main → limelight (Rush), v3 → error (The Warning),
v4 → rivendell (Lord of the Rings), testing stays as testing.
Suite field unchanged — codename is an alias, both work in sources.list.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Multipath tested and confirmed working on TrueNAS CORE 13.x (FreeBSD/ALUA)
and SCALE 24.10 + 25.04 (Linux LIO). Failover and path recovery validated
on both platforms. Closes#256.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
If /etc/multipath.conf already had a manual devices{} block before package
install, postinst would append a second identical block, causing multipathd
to emit "duplicate keyword: devices" warnings. Strip any unmanaged devices
block in the append path to keep the file clean.
Fixes#279. Discovered 2026-06-20 on pve01-hq during SCALE 25.04 multipath testing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Old multipath testing debs were never pruned, causing apt to see multiple
3.1.0~beta+<sha> versions and pick whichever sha sorts highest — not newest.
Mirror the existing pool/testing cleanup that already did this correctly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Without a registered JS panel the truenas-multipath type was invisible in
the PVE Datacenter → Storage → Add dropdown. Adds truenas-multipath.js
with all base TrueNAS fields plus the required truenas_portals field, and
wires it into the multipath package install/remove scripts and CI build.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
On install: if /etc/multipath.conf already exists, append a tagged
block rather than skipping silently (old) or overwriting (dangerous).
On upgrade: replace only the tagged block, leaving the rest untouched.
On remove/purge: strip the tagged block; rest of the file preserved.
This means admins who added their own stanzas during manual testing
(e.g. before the package existed) won't lose config on install or
upgrade. Closes the silent "file already exists" footgun.
Also fixes pre-existing SC2015 shellcheck warnings in postrm.
Refs #256
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
TrueNAS.pm::qemu_blockdev_options returns an iscsi:// blockdev — wrong for
multipath where path() returns /dev/mapper/<wwid>. Simply removing the override
caused the parent's iscsi driver to be inherited instead.
Fix: explicitly call PVE::Storage::Plugin::qemu_blockdev_options (the
grandparent). It calls path(), sees /dev/mapper/<wwid> starts with '/',
stats it as S_ISBLK, and returns { driver => 'host_device', filename => ... }.
QEMU receives: "driver":"host_device","filename":"/dev/mapper/<wwid>"
Tested: VM 103 started with multipath disk (vm-103-disk-0 on TrueNAS-Scale01-MP).
QEMU confirmed using host_device driver on /dev/mapper/36589cfc000000ad...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
TrueNAS.pm: expose naa field from _find_extent return hash — needed by
TrueNASMultipath.pm to build the /dev/mapper/<wwid> path.
TrueNASMultipath.pm: use multipathd reconfigure + fallback multipath -v0
instead of bare multipath(8) which conflicts with a running multipathd
daemon. Tested end-to-end on pve01-hq against TrueNAS SCALE:
- Two sessions (172.31.69.91 + 192.168.69.91) both established
- /dev/mapper/36589cfc000000eb2adcd24154021d376 appears as block device
- deactivate_volume cleanly flushes and logs out
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>