Commit Graph
104 Commits
Author SHA1 Message Date
Nikolay MiroshnichenkoandClaude Opus 5 ec63bcf106 load_ga_export, entity naming, climate setpoint shift; fix suggest_repairs crash; CI runs every test
Three things found by auditing a TapPlan export, plus one regression found on the way.

load_ga_export(path): new tool. Reads an ETS ga-export/01 XML (ETS "Export Group
Addresses", or the import file a planning tool writes) into a project without
devices, through safe_fromstring, capped at 50 MB and 8 levels of range nesting.
Invalid or duplicate addresses and unknown DPT tokens go to import_warnings, never
silently. Round trip with our own generate_ets_xml is covered by a test.

Entity naming: lights, covers and climates are named after the common word prefix
of their member names, cutting only function words. The first version of the rule
turned "01. <room> - All Blinds - Move" into "01" on a real project, so anything not
in the function vocabulary now stays. On six real projects 214 of 1189 entities got
a shorter name, each rename reviewed, no address mapping changed. Regenerated
packages show different names; noted in the changelog.

Setpoint shift: 9.002 / 6.010 with a shift word maps to setpoint_shift_address,
setpoint_shift_state_address and setpoint_shift_mode (keys checked against the HA
KNX climate docs). It used to become a plain sensor.

Fixed: suggest_repairs raised UnboundLocalError on any project with a missing status
GA. My local rename in the 12.09 typing cleanup left two references to the old name.
test_council_fixes covers it but was not in CI, so it shipped. CI now runs all 21
test files instead of a hand-picked 10; the 11 added ones pass from a clean clone.

Verified: all 21 tests, ruff, mypy with the package installed, corpus guard no drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 00:11:45 +02:00
Nikolay MiroshnichenkoandClaude Opus 5 29bfba7684 tools: corpus guard counted missing_value_status twice
detect_missing_status already folds in detect_role_completeness, and the guard
called both, so every missing_value_status finding was recorded twice. The
product tools (server, report, handover) only call the aggregate and were never
affected. Found while auditing a TapPlan export with the same call pattern.

Baseline re-recorded in this commit. Exactly 12 metrics moved, all of them that
one code and the warning total it feeds, halved on each of the six projects;
nothing else changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:57:55 +02:00
Nikolay Miroshnichenko e4f0f00c45 ci: restrict GITHUB_TOKEN to contents: read
First CodeQL run flagged the workflow for running on the default token
permissions. It only checks out and runs tests, so read is all it needs.
2026-09-12 21:42:03 +02:00
dependabot[bot]andNikolay 44a2296d3a ci: bump the actions group with 2 updates (#14)
Bumps the actions group with 2 updates: [actions/checkout](https://github.com/actions/checkout) and [actions/setup-python](https://github.com/actions/setup-python).


Updates `actions/checkout` from 4 to 7
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v4...v7)

Updates `actions/setup-python` from 5 to 7
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v5...v7)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
- dependency-name: actions/setup-python
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Nikolay <nickol@me.com>
2026-09-12 21:41:44 +02:00
Nikolay Miroshnichenko 39be07ebd6 ci: keep Dependabot off mcp major bumps until the v2 port
mcp 2.x drops mcp.server.fastmcp; the import fails outright on 2.2.0, which is
what Dependabot's first PR would have allowed. Minor and patch updates still flow.
2026-09-12 21:40:53 +02:00
Nikolay Miroshnichenko af6d0e24c0 ci: run mypy against the installed package, fix what that exposed
The lint job failed on its first run with three errors my local run had not
shown. Cause: locally mypy had no xknxproject installed, so ignore_missing_imports
hid every mismatch against that library's own TypedDicts. CI installs the package
and therefore sees them.

The two range-walk helpers took dict[str, Any] but are handed a GroupRange
TypedDict, which mypy rejects because dict is invariant; they take Mapping now,
which is what they actually use. suggest passes the raw parsed dict into
build_loaded_from_raw, whose parameter is KNXProject, so the boundary is an
explicit cast rather than a widened public signature.

CONTRIBUTING now gives the command that matches CI, with the reason, so the next
person does not get a clean local run and a red CI.

Suite passes, corpus guard reports no drift.
2026-09-12 21:38:54 +02:00
Nikolay Miroshnichenko 264767f91f ci: ruff + mypy gates, Python 3.13/3.14 in the matrix, Dependabot
The CI ran tests only. Added a lint job with two narrow gates:

ruff with select = E9,F — syntax errors and pyflakes (undefined names, unused
imports, broken f-strings). Style rules are deliberately off: the full default
set flagged 202 things, almost all cosmetic, and a gate nobody reads is not a
gate. Fixed the 9 findings it did have (8 unused imports, 1 empty f-string).

mypy, clean at 13 errors fixed. None of the 13 was a live bug, so every fix is
an annotation or a local rename: the two accumulators in advanced, the block map
in appprog_parser, a reused loop name in project and in repair, a TypedDict lost
through dict(), a mixed-type summary dict in handover. Two are worth naming.
_CH_TOKEN_RE now captures the letter prefix directly instead of re-matching its
own capture, which removes an unguarded .group() on a possibly-None match; the
regex accepts exactly the same strings. repair's status-DPT lookup now spells out
the None case, which is what dict.get already did since None is not a key there.

Verified the fixes moved nothing: the full suite passes, and the corpus guard
reports no drift across all six real projects, the 3646-GA one included.

Matrix extended to 3.13 and 3.14; both were smoke-tested locally first.
Dependabot config for pip and github-actions, weekly and grouped so a quiet week
is one PR. Dependabot security updates and CodeQL default setup enabled in the
repository settings.
2026-09-12 21:36:23 +02:00
Nikolay Miroshnichenko 4eb54da53c tools: local corpus guard against real ETS projects
Public CI only runs synthetic fixtures, but every regression we have actually
shipped showed up on real projects, which are confidential and cannot reach a
GitHub runner. corpus_check.py runs a local corpus through the full pipeline
(parse, checks, HA YAML, entity suggestions) and diffs the counts against a
recorded baseline.

Only numbers are committed: counts, severity and code histograms, and a short
digest per source file. The project map with real paths lives in
tools/corpus_map.json, which is gitignored; corpus_map.example.json shows the
shape. Verified both directions: a deliberate change to the diagnostics filter
made the guard exit 1 and name the two metrics that moved, a clean tree exits 0,
and a machine without a map exits 2 without doing anything.

Wired into CONTRIBUTING and into the release checklist before the build step.
2026-09-12 21:31:44 +02:00
Nikolay Miroshnichenko 4010743eb5 tools: stable ordering + cursor paging for list_group_addresses and get_devices
- sort by the address itself (GA main/middle/sub numerically, device area/line/device), so the
  result no longer depends on parser dict order and a retry returns the same page
- return {items, total_matched, returned, next_cursor} instead of a bare list truncated at limit:
  truncation was silent and a caller could not distinguish a full answer from a cut one
- tests/test_paging.py (in CI): numeric vs lexical order, full walk at page sizes 1..100 covers every
  row exactly once, same page after a reshuffled re-parse, filters reflected in total_matched
- verified on a real 3646-GA project; README EN/RU + CHANGELOG updated
- prompted by feedback on the r/mcp post (2026-09-12)
2026-09-12 21:15:55 +02:00
Nikolay Miroshnichenko 07352cdb61 suggest: split multi-output channels by vendor object marker; drop device diagnostics
- _subunit()/_split_subunits(): '[1] Switch On/Off' style markers separate the outputs a vendor
  packs into one ETS channel; each sub-unit is classified on its own (ids and secondary_info carry it)
- diagnostics (error/communication/firmware/version/bus voltage/reset, EN+DE+RU) never become sensors
  and are excluded from the name-based fallback; counted in hints.diagnostics_skipped
- real-project effect (3646 GAs): entities 341->513, sensors 752->367, platform agreement 92->95 %,
  address-key agreement 94->95 %
- tests: multi-output split (two lights, statuses consumed, no sensor fallout) + diagnostics filter
2026-09-12 09:09:08 +02:00
Nikolay Miroshnichenko 80fc279ac1 docs: release checklist (PyPI + official MCP Registry), with the traps hit during 0.8.1 2026-09-12 09:02:08 +02:00
Nikolay Miroshnichenko 1ac0527c2c 0.8.1: published to PyPI and to the official MCP Registry
- server.json name fixed to io.github.NickoScope/nickol-knx-mcp (registry permissions are case-sensitive)
- README mcp-name marker in the same case; 0.8.0 had it lowercase, which the registry rejects
- version bumped in pyproject + __init__; CHANGELOG 0.8.1
- registry: io.github.NickoScope/nickol-knx-mcp v0.8.1 active, package pypi/nickol-knx-mcp 0.8.1
v0.8.1
2026-09-12 09:00:12 +02:00
Nikolay Miroshnichenko 9fe4df1dfc server.json: shorten description to registry limit (<=100 chars); validated OK 2026-09-12 08:54:37 +02:00
Nikolay Miroshnichenko f83819c434 changelog: merge duplicated Added heading 2026-09-12 08:41:24 +02:00
Nikolay Miroshnichenko 3d134b3bc7 registry: server.json + mcp-name ownership marker in README (official MCP Registry prerequisites)
- server.json: schema 2025-12-11, registryType pypi, name io.github.nickoscope/nickol-knx-mcp,
  stdio transport, NICKOL_KNX_WORKSPACE / NICKOL_KNX_CATALOG documented as optional env vars
- README: <!-- mcp-name: io.github.nickoscope/nickol-knx-mcp --> (becomes the PyPI description,
  which is what the registry checks for ownership)
- verified: uv build produces sdist+wheel, twine check PASSED, marker present in wheel METADATA
2026-09-12 08:40:55 +02:00
Nikolay Miroshnichenko aea59066ba suggest.py: structure-first entity suggestions prototype (HA SuggestionProvider #2 contract)
- channel -> object flags (write=sink, transmit=source, dual ignored) -> DPT pattern
  (cover, climate, light/switch, sensors) -> vendor object texts -> names as tie-break
- FB-covered channels skipped (also in fallback); shared status GAs excluded;
  duplicate writers deduped; pseudo-channels for channel-less devices; unwired flagged
- tests/test_suggest.py: FB-provider nameless fixture reproduced from structure (CI)
- tools/eval_suggest.py: agreement vs name/Function engine on real projects
- CHANGELOG (Unreleased/Added, experimental)
2026-09-10 12:18:49 +02:00
Nikolay Miroshnichenko a32b1970f3 check_policy: example profile seeded from the loaded project's own main groups (issue #13)
- example_policy_yaml(project): mains written with inferred domain, range name and
  category mix; no-majority mains commented-out with their mix; defaults only as a
  labelled comment; reserve.expect_range follows the project; empty/2-level -> {}.
- taxonomy_seed() is the single source for _infer_taxonomy and the example.
- server.check_policy(write_example_to) passes the loaded project, reports seeded_from.
- tests/test_policy.py: cases 4-6 (no leaked mains, round-trip, mixed main, quotes, no project); added to CI.
- README/README.ru/CHANGELOG.
2026-09-09 20:03:32 +02:00
Nikolay MiroshnichenkoandClaude Fable 5.1 05a0658a3a ux: gentle real-project feedback nudge where people actually run the tool
Cloners install and run the tool but rarely reread the README, so the existing
'call for testers' never reaches them. Two low-noise touchpoints instead:
- project_report: a one-line footer linking the real-project test-report issue
  template (the report is the artefact a human reads after a real run);
- load_project: a short 'feedback' field on the result — the first call every
  new user makes.
Anonymised addresses explicitly welcome. No logic change; full suite green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 12:19:20 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 467772618e fix(#12 residual): recognise a shutter Function by its 1.008 move, not only a role token
Kris1166's real-file re-test of f20b2e4: main bug gone (light:0, all 29 RM → status) but
15/29 position feedbacks stayed lighting/not_mapped. Root cause (reproduced 1:1): those
actuator-level Functions carry blank or raw-GUID roles — no role token anywhere — so the
shutter-Function gate skipped them even though they hold a 1.008 up/down move. The
anomaly (2/3/1 promoting while its step/stop stayed unknown) is explained: its move has a
role; its step/stop is DPT 1.009 'Enable' (the reporter's own modeling slip, not a step).

- A shutter Function is now recognised by a shutter ROLE token OR by an UNCONTESTED 1.008
  move member (category==shutter; a 1.008 whose name pins another domain must not license
  its neutral siblings — gate-1 audit).
- is_step includes DPT 1.017 (trigger), which the cover builder already treats as a stop.
Gate-1: APPROVED; 0 category/kind/platform drift across 5 real .knxproj (1965 GA); the new
fixtures fail on the old code. Gate-2: the LLM council already endorsed the sibling-DPT
(1.008 move) signal as the primary structural evidence — this implements it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-04 21:33:25 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 f20b2e4126 fix(#12): recognize a '… RM' (Rückmeldung) shutter feedback as the cover's position status, not a light
Field data by Kris1166 (issue #12). A DPT-5.001 shutter position feedback named only '… RM'
was classified lighting/light and emitted as a bogus HA light entity. Two layers:
- classification (project.py): an ETS Function carrying a shutter role promotes a member to
  shutter — but only on POSITIVE evidence (member is a step/stop 1.007/1.010 OR a 5.001 position
  STATUS, the Function has an up/down move sibling, and the name has no explicit non-shutter
  domain), so a scene-recall/dimmer-feedback/lock sharing a mixed Function is left alone
  (LLM-council review; a Function is a free grouping, not proof of domain).
- status recognition (project.py name_is_status): a standalone 'RM'/'FB' token is now a status —
  the 'rm '/'rm_' substrings and the stopword tokenizer both missed a name ending in bare 'RM'.
- analyze._is_status_ga delegates to name_is_status; generate_ha then wires it as the cover's
  position_state_address.
Passed both gates (self-audit + LLM-council). tests/test_rm_status.py pins it (red/green + negatives:
a light's RM stays light, a stray light / a scene-recall in a shutter Function are not promoted).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-02 10:02:06 +02:00
Nikolay Miroshnichenko fb53018b59 docs(readme): align spec→design headline with the case study (96%, 662 vs 687 GA) in EN+RU — was an inconsistent ~92%/3,600-GA figure vs the 96% badge 2026-08-29 15:30:16 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 5c87169129 harden(#11): operation!=domain gate, deterministic tiebreak, no central-stop steal (LLM-council P0 + gate-1)
Post-fix hardening from an LLM-council review of 0198ae3, with a gate-1 audit pass:
- _is_shutter_control rejects a bare 1.007/1.010 whose tokens are a foreign domain
  (light/dim/socket/HVAC) or a central macro (Alle/Zentral) — a 'stop' is an operation
  signal, not a domain one, so a lighting 'Licht Stopp' or a house-wide 'Alle Stopp' can
  no longer be mis-paired/stolen as a cover's step/stop.
- _rank gains a stable address tiebreak -> deterministic pick across parses on a full tie.
- gate-1 caught 'master' wrongly excluding Master-zone covers; removed. Vent words kept
  admissible (roof-window/flap openers are covers).
- tests groups E/F/G pin the foreign-stop, central-steal, and Master-zone cases.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-29 15:01:25 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 0198ae3019 fix(#11): pair every cover to its own step/stop across all 3 real ETS Function structures
Second pass on issue #11, grounded in Kris1166's real ETS-6 field dump (groups A/B/C),
after the first fix (dd8ecac) proved partial. Three levers:

- project.py: ETS Function role (MoveUpDown/StopStepUpDown) authoritatively promotes a
  GA to category=shutter — rescues bare DPT 1.007 step/stops the name classifier left
  'unknown' (group A: 26/33 covers were dropped). Guarded to DPT 1.x/5.x; role strings
  from Kris's real dump, not invented.
- generate_ha.py: cover builder now gathers candidates and picks the BEST per role
  (same-function > name-token overlap > nearest sub) instead of first-match, so a
  function-less roller takes its own step/stop, not an awning's sharing the zone token
  (group C). Bare 1.007/1.010 admitted only with a shutter name signal, so a foreign
  lighting step can't be stolen (audit finding).
- tests: test_cover_pairing.py rebuilt from the real A/B/C structures + a group-D guard
  for the foreign-1.007 case. Fails without the fix, passes with it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-29 13:13:59 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 dd8ecac090 fix(ha): pair a cover's step/stop within its ETS Function, never across it (issue #11)
generate_ha_package cross-wired move_short_address across unrelated ETS
Functions sharing a zone token (a window and an awning both 'Room A'), and
could leave a cover with no step/stop. Pair siblings by ETS Function
membership (authoritative); a sibling owned by a different function is never
borrowed. Falls back to the existing name-token logic only when the move GA
has no function, so projects without ETS Functions are unchanged.

Adds tests/test_cover_pairing.py (reproduces the field case from #11, fails
without the fix) and wires it into CI.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-27 11:01:54 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 b38c25675a docs(roadmap): announce upcoming Logic Machine support (in research)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-18 17:58:08 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 c1e9d206c5 feat: fold topology findings into project_report (section 2.4 + totals)
detect_topology_issues now surfaces in the human-review project_report markdown
(new '2.4 Topology & addressing' section) and its severities count into the
report totals — closes the audit's deferred LOW (per CLAUDE.md the human reviews
project_report before ETS import, so topology issues must be visible there).
Functional check: Poltavskiy report renders the coupler INFO; suite 27 passed
(pre-existing test_protocol collection error untouched).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 22:14:11 +02:00
Nikolay MiroshnichenkoandClaude Opus 4.8 8312d531eb feat: check_topology — per-line device count + individual-address validity (KNX canon)
New design-time check grounded in the KNX standard (KNX Handbook + xknx via
deepwiki). detect_topology_issues() + check_topology() MCP tool (30->31 tools),
folded into analyze_all counters:
- devices per line vs TP1 segment 64 / line 256 (KNX Handbook p.36/40/55) — TP-only
- individual address valid A.L.D (area 0-15, line 0-15, device 0-255; xknx-confirmed)
- duplicate individual address -> error; missing .0 coupler on a multi-line TP -> info
Medium-gated on the real xknxproject string ('Twisted Pair (TP)'), so IP/PL/RF
lines are never falsely flagged (fixed a dead 'medium==IP' guard caught in audit;
regression test added). senior-code-audit: APPROVED. Functional smoke on 2 real
projects: IP lines clean, TP coupler/count notes correct. 64/256 cited to the
Handbook, not attributed to xknx. tests/test_topology.py 7/7; suite 27 passed
(pre-existing unrelated test_protocol collection error untouched).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 21:41:14 +02:00
Nikolay1andClaude Opus 4.8 bf7eba8898 room-library: Theben-benchmark corrections (dogfood on 5 official ETS examples)
Ran KNX Association-ecosystem vendor (Theben) 5 official ETS example .knxproj
through our QA. Findings folded back into the product:

- repair.py _infer_dpt: relative dimming (Brighter/darker, heller/dunkler,
  светлее/темнее) now infers 3.007 DPT_Control_Dimming instead of falling to
  1.001; explicit '% / value' stays 5.001; scene checked before the broad
  value branch. DPTs verified against XKNX/xknx via deepwiki.
- room_library.py: new reusable blocks distilled from vendor patterns —
  shutter_venetian (slat + slat-status + wind/frost safety-lock), climate_floor
  modulating actuating-value + status, ventilation (airing stage + CO2 +
  humidity), lighting_dali, multisensor_air, central scene/all-off/all-on.
- project.py classifier: airing/Lüftung/проветривание -> hvac (was unknown).
- analyze.py: safety_input_no_status INFO exemption for write-only wind/frost
  lock inputs (Theben central wind-safety pattern).
- SCHEMA.md: new function rows; automation_intents implementation: knx|external
  with safety_related => knx rule; DPT-tolerance note (missing-DPT hard error
  applies to generated output; ETS4 object-derived DPT on import is legitimate).
- templates: ventilation slot (bathroom/kitchen), opt-in venetian + wind/frost
  safety intent (living/bedroom), new house-level central.yaml.
- tests: +test_theben_corrections.py (10); golden bedroom 26->28 GA (climate
  actuating-value). pytest 20 passed (pre-existing unrelated test_protocol
  collection error untouched). senior-code-audit: APPROVED.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 11:21:18 +02:00
Nikolay1andClaude Opus 4.8 f0d1afdce2 docs(readme.ru): bring RU tool table to parity with EN (25 → 30)
Add the 5 tools missing from the RU table (present in EN + server.py):
- explain_ga            → Чтение
- check_policy          → Валидация
- check_device_parameters → Починка и дизайн
- validate_room_template, compose_rooms → new Room Library group
Descriptions mirror EN meaning, signatures verified against server.py.
Bump counters: '## 5. Инструменты MCP (25)' → (30) and package layout
'25 инструментов' → 30. RU table now 30 rows, 1:1 with EN.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 22:09:22 +02:00
Nikolay1andClaude Opus 4.8 ce64250eac docs(readme): align narrative with product (Room Library R1 + shipped roadmap)
Positioning-only sync (tool table untouched):
- Intro: 'Three things' → 'Four things' + new bullet 'Compose a new project
  from parametrised room templates' (dry-run, new projects only, R1).
- Add Scenario 4 (compose_rooms → allocation manifest + ETS XML/CSV + device
  BOM; generated .knxproj re-read by standard loader, 0 errors / 0 warnings;
  R2 docking + device selection flagged as planned).
- Roadmap: reframe check_device_parameters and check_policy as shipped (they
  are already in the tool table); add Room Library R1 shipped / R2 planned;
  keep ETS7 Smart Linking note.
- Package layout: add room_library.py + room_templates/.
- README.ru.md: mirror the intro, Scenario 4 and package-layout edits (ru has
  no Roadmap section). NOTE: ru tool table still lags at 25 vs EN 30.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 22:04:53 +02:00
Nikolay1andClaude Opus 4.8 3d0aa8ec29 docs(site): sync landing to 30 tools + add Room Library scenario
Tool count 25 → 30 (meta description + tools card). Add a 4th scenario
card 'Compose a new project' (compose_rooms → allocation manifest + ETS
XML/CSV + device BOM; dry-run, new projects only, R1). Clarify the
read-only thesis: never mutates your source .knxproj or touches the bus.
Update 'three'→'four scenarios' heading + anchor. R2 (docking + device
selection) flagged as planned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:55:00 +02:00
Nikolay1andClaude Opus 4.8 377cba6e30 Merge feat/room-library-r1: Room Template Library R1 (28→30 tools)
Compose a new KNX project from parametrised room templates. Adds two MCP
tools (validate_room_template, compose_rooms), six built-in room templates
with a public SCHEMA.md contract, a resolved-IR allocator, and a round-trip
.knxproj generator re-read by the standard loader. Generated house passes
all four linters with 0 errors / 0 warnings. 14/14 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:47:49 +02:00
Nikolay1andClaude Opus 4.8 1a44f348be room-library R1: docs + packaging (README 28->30, CHANGELOG, package-data)
Document both new tools in README (tool count 28 -> 30), add a CHANGELOG entry
under Unreleased, and ship the room_templates/*.yaml + SCHEMA.md as package data.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:26:23 +02:00
Nikolay1andClaude Opus 4.8 9b337a4882 room-library R1: tests (golden RU/EN, idempotency, invariance, round-trip, oracle)
tests/test_room_library.py covers all R1-required cases: exact golden RU/EN
object maps for one template; idempotency; room permutation invariance;
sub-address exhaustion raising a clear error; round-trip (generated house ->
load_project -> 0 errors AND 0 warnings across validate_naming /
detect_missing_status / detect_dpt_issues / check_policy); a hand-written
negative-oracle manifest vs fact; plus per-slot preset override, template
rejection and BOM proposal checks.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:26:23 +02:00
Nikolay1andClaude Opus 4.8 bdb93cd128 room-library R1: register validate_room_template + compose_rooms MCP tools
Two new tools (28 -> 30): validate_room_template checks a built-in or custom
template against the R1 schema; compose_rooms builds a new project from a room
list and returns the allocation manifest, ETS GA XML/CSV (via the existing
generators on the re-read project), and a device BOM proposal. New projects
only, dry-run by default; writes artifacts into the confined workspace when
dry_run=false and output_dir is given. Never touches a bus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:26:23 +02:00
Nikolay1andClaude Opus 4.8 ebc03f7118 room-library R1: template format + resolver/IR + composer core
Add the Room Template Library core: the public YAML template contract
(schema_version, locale-neutral semantic slot_id, ru/en labels, per-slot
basic/comfort presets, area_m2 as a provenance-tagged hint) with 6 built-in
rooms (bedroom, children, living, kitchen, bathroom, corridor) + SCHEMA.md.

room_library.py implements a function-first pipeline: templates+params ->
resolved IR dataclasses (NOT GARecord) -> deterministic, permutation-invariant
allocation (main=domain, middle=role, sub sequential; hard error on sub
exhaustion) -> a real .knxproj ZIP that is re-read via load_project (same path
as third-party projects, so generation never touches the classifier). Includes
validate_room_template, compose(), manifest + device BOM proposal builders.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:26:11 +02:00
Nikolay1andClaude Opus 4.8 4a9dd5ce4a fix(generate_ha): Russian function words in identity + exact-identity outranks subset
Dogfooding the voice layer exposed two pairing defects: (1) _FUNC_WORDS had no
Russian entries, so 'вкл/выкл/цвет/яркость' polluted device identity and an RGBW
colour GA could not find its own switch; (2) subset identity matching let
'RGBW подсветка яркость' {zone, подсветка} grab 'камин подсветка' by subset and
its status by tie-break. Sibling pick and status_for_dpt now prefer EXACT
identity equality before subset fallback. 13/13 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 12:31:37 +02:00
Nikolay1andClaude Opus 4.8 6cabea08af fix(P5/policy): dogfood round 3 — location stopwords, greedy-term pruning, reserve check
- location phrases no longer vote a domain: "Boiler room light" is a light in
  the boiler room, not an HVAC function (location stoplist stripped first).
- dropped the bare "температур" hvac term: a temperature GA takes its domain
  from context (уставк/кондиц/тёпл words or its main group), so "Гостиная
  температура" in a Sensors main is a sensor, not an hvac actuator outlier.
- "online/offline/heartbeat/watchdog/связь" are diagnostics terms (checked
  first), so actuator-monitoring GAs named after the load they watch stay diag.
- check_policy: central macros ("Весь свет выкл") are never taxonomy outliers;
  policy_no_reserve now also fires under an INFERRED taxonomy when no reserve
  main exists at all; "весь "/"мастер" added to central-macro tokens.

Messy-house dogfood now reports exactly the planted misplacements (3/3) and the
reserve gap; demo 0 outliers. 13/13 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 11:36:50 +02:00
Nikolay1andClaude Opus 4.8 ccfedbdeca fix(P5): dogfood round 2 — presence is a sensor, 9.001 soft, passive-range rule
- DPT 1.018 occupancy: category sensor (was diagnostics; the name always says
  sensor and the table fought it into a fake conflict).
- 9.001 temperature is SOFT: room temperature defaults to hvac, but a weather-
  named temperature re-domains to sensor without a fabricated conflict.
- illuminance terms: "освещённость"/"illuminance" are sensor words and must not
  be swallowed by the greedy lighting prefix "освещ".
- passive-range rule: a measurement inside a Sensors/Energy/Diagnostics main
  takes that main's domain; a COMMAND is never retyped by a passive range.

Deliberately NOT tuned further: remaining real-project findings (Minsk 10/685,
Razdory 75/3646, demo 2/239) are defensible mixed-main observations a reviewer
should see once — zeroing them out would overfit one integrator's idiosyncrasy.
13/13 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 11:33:20 +02:00
Nikolay1andClaude Opus 4.8 3c086fe228 fix(P5): classifier dogfood round — range-map depth bug, fallback-DPT softness, RU terms
Building a deliberately messy synthetic house instantly exposed four real gaps:

- _build_range_name_map: a middle group 0 starts at the same raw address as its
  main, so the modulo heuristic let "Вкл/Выкл" clobber "Освещение" and the
  main-range rescue silently died. Depth in the range tree decides now.
- a category from a whole-main DPT fallback (e.g. 20.609 -> hvac) was treated as
  strong; only an exact (main,sub) table entry is strong now (is_exact_dpt).
- DPT 1.010 start/stop marked soft: ventilation timers use it, not just shutters.
- terms learned: Cyrillic а/с, сплит, вытяжк, яркост (RU only); weather/метео in
  sensor, and sensor now outranks hvac so weather-station temperature is a
  sensor, not an HVAC actuator.

Real-project noise after the round: Minsk 1 outlier/685 GAs (was 31 pre-P5),
demo 2/239, Razdory 46/3646. Regressions in test_domain_classifier. 13/13 green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 11:27:22 +02:00
Nikolay1andClaude Opus 4.8 e9105d45aa feat(P5): contextual domain classification — DPT is one signal, not the verdict
A 1-bit switch is domain-agnostic, yet DPT 1.001 defaulted to 'lighting', so an
"AC on/off" was classified lighting and tripped downstream checks. The domain is
now a combination: explicit name > strong domain-encoding DPT > main-group name;
a bare 1-bit GA with no signal is honestly 'unknown', never a guessed lighting.
An explicit name contradicting a STRONG DPT yields 'unknown' (a genuine conflict,
shown by explain_ga as contested), not a silent pick. Word-boundary matching so
"ac" no longer fires inside "terrace"; name terms broadened (EN+RU HVAC incl.
ac/a-c/air-conditioner, RGB/colour lighting).

check_policy taxonomy-outlier now flags only misplaced ACTUATORS (lighting/
shutter/hvac); sensors, scene links, thresholds and timeout params are the
cross-cutting GAs they are. HA generator still emits a switch for a now-HVAC
simple on/off, so no device is dropped.

Verified on 3 real projects: every AC on/off is HVAC; outlier noise fell demo
11->2, Minsk 31->6, Razdory 200->57. tests/test_domain_classifier.py (13/13 green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 10:44:04 +02:00
Nikolay1andClaude Opus 4.8 1a60264c78 feat(P7b): explainable aggregate scores — expose denominator + formula
Matter readiness, the completeness grade and command/status coverage each now
carry a `math` block: numerator, denominator, the exact formula that reproduces
the percentage, and (for Matter) the functions excluded from the denominator
because their category has no Matter cluster — previously a silent skip and the
biggest "why is this number what it is?" gap. Completeness also states its band
thresholds. Report-only, additive; scores unchanged, only now auditable.
Raised by external review. tests/test_explainable_aggregates.py (12/12 green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 10:15:40 +02:00
Nikolay1andClaude Opus 4.8 f68c85bb48 feat(P7a): harden parsing of untrusted .knxproj (zip-bomb / XXE defense)
A .knxproj is a user-supplied ZIP-of-XML — untrusted input. New safexml.py
centralises hardened parsing for every place we open an archive or parse XML
ourselves (load_project, check_device_parameters, app-program parser):

  XML  - dependency-free reject of DOCTYPE/ENTITY (billion-laughs / XXE;
         legitimate ETS XML never carries one) + defusedxml parser-level
         blocking when installed (added as a dependency).
  ZIP  - pre-flight against absolute caps (archive size, entry count, per-member
         and total decompressed size, compression ratio) rejects zip-bombs;
         member names checked for path-traversal / absolute paths; each member
         read through a streaming cap so a lying header can't exhaust memory.

Violations return a normal error dict, not a traceback. Still read-only.
Raised by external security review. tests/test_safexml.py (11/11 suite green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 10:11:38 +02:00
Nikolay1andClaude Opus 4.8 84c1c7dbd9 fix(P3): root-cause suppression — an empty-name GA no longer cascades
Council #2: the empty GA 2/5/2 spawned missing_status + policy_taxonomy_outlier +
a lighting classification on top of its real defect. An empty name means the
classifier can't be trusted, so detect_missing_status and check_policy now skip
blank-name GAs; the single root cause is empty_name (check_naming). Demo 2/5/2:
3 findings -> 1. tests/test_root_cause.py; full suite green (10 tests).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:29:54 +02:00
Nikolay1andClaude Opus 4.8 10169275e3 feat(P1): explain_ga — provenance/confidence for one GA (27->28 tools)
Three independent reviewers (two external councils + a field integrator) asked for
the same thing: the enriched model mixes ETS facts, DPT structure and name
heuristics, and downstream tools treat it almost as fact. explain_ga makes the
reasoning auditable per GA — evidence per decision with a confidence tier
(authoritative ETS Function > structural DPT > heuristic name), the status-pairing
method, and CONFLICTS (name 'AC' vs DPT 'lighting' -> contested), the silent-
misclassification hotspot. Additive, read-only, no core-model change. Full suite green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:27:10 +02:00
Nikolay1andClaude Opus 4.8 5036a19b4b feat(P2): role-aware feedback completeness — catch a value command with no value status
'Does the function have *a* status' missed a dimmer with on/off status but no
brightness status (demo planted error #3), silently inflating coverage/Matter.
detect_role_completeness flags a brightness/position command (5.001) whose device
has no matching value status (missing_value_status), with a device-identity match
so it doesn't borrow a sibling's status. Demo recall 4/5 -> 5/5. Raised by the
external expert review. Test + demo ground-truth updated. Full suite green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:23:02 +02:00
Nikolay1andClaude Opus 4.8 c53f388dcb fix: council-review noise cuts — scenes/central-macros exempt, language-matched repairs
External expert review (real run on the demo house) flagged: scenes (18.001)
wrongly marked missing_status + an unsafe DPT-1.011 status repair for them;
'All blinds down' warned while 'All lights off' was correctly INFO; and a Russian
'(статус)' suffix synthesised on an English project. Fixes: scene_no_status INFO
+ no synth status for scenes; broadened central-macro detection (generic 'all
blinds/shutters/...'); language-matched repair suffixes. Demo missing_status
warnings ~7 -> 1 genuine. tests/test_council_fixes.py; full suite green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:09:57 +02:00
Nikolay1andClaude Opus 4.8 1e924ede6f feat(param_check): significance layer — a focus list of config-value outliers
Rank findings by significance: config_value (setpoint/hysteresis/time/threshold —
an odd device is usually a real mistake) vs config_flag (mode/type) vs label
(per-room text, often intentional). New focus list surfaces the config-value clear
outliers first. On a real 275-device project this turns 422 raw outliers into a
9-item focus that includes the thermostats whose init-setpoint differs from their
siblings. Multilingual name hints (EN/DE/RU). Test updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:28:32 +02:00
Nikolay1andClaude Opus 4.8 81799169c5 feat: check_policy — Project Policy Profile (26->27 tools)
Validate a project against ITS OWN rules, not a universal standard. With a YAML
profile the declared main-group taxonomy / naming / pairing is authoritative;
with no profile the taxonomy is INFERRED from the project itself and GAs that
deviate from their main group's own majority domain are flagged — never against
an alien standard (a well-organised real 356-GA project drops from 329 false
mismatches vs the default to 31 genuine self-deviations). Answers the recurring
integrator critique that 'your best practices aren't universal'. Example profile
+ test_policy.py included.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:06:59 +02:00
Nikolay1andClaude Opus 4.8 4c36bc182e feat: check_device_parameters — cross-device parameter outlier QA (25->26 tools)
New module param_check.py + MCP tool check_device_parameters: reads per-device
ParameterInstanceRef values from the .knxproj project part (xknxproject does not
expose them), groups identical devices by application program, and flags the odd
one out — clear_outliers (a strong majority with a small minority, e.g. one
thermostat with a different setpoint/hysteresis) and split_configs (balanced
variants → review). Parameter names resolved from the app-program; encrypted
projects skipped honestly. Read-only, no ETS/bus. Validated on real 42-275-device
projects; synthetic test_param_check.py. Answers a community feature request.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 21:53:32 +02:00