Upgrading HotLoop Gateway to 4.17.0
Two ways onto 4.17.0, and they are nothing alike. From 4.16.0 it’s a plain helm upgrade with one flag that bites, and afterwards a new Secret is as much a part of your plant’s data as the database is. From 4.3.1 it’s a renamed chart that a plain helm upgrade can’t touch, and you go through 4.16.0 on the way. Find the one that’s yours and read all of it before you touch a running plant.
Upgrading from 4.16.0
Section titled “Upgrading from 4.16.0”Why bother: 4.16.0 has bugs a running plant can trip over, and 4.17.0 fixes them. The command prompt can swallow a Set, so an operator watches the prompt close, figures the setpoint went out, and nothing was written. Saving a rule in the editor, even to rename it, quietly resets its run budget to five minutes. A device file pushed to an Edge Relay never lands, while the Gateway says config pushed. And an EmberNET App Store deploy loses its tenant. On top of the fixes you get recipes, scripts, blueprints, helpers, the logbook, the new automation steps, read-only UniFi, and MCP going from 16 tools to 25.
It’s a plain helm upgrade. We rendered both charts at both versions: every selector and every volume claim is identical, and the only new object is the key’s Secret.
helm repo update hotloophelm upgrade <release> hotloop/hotloop --version 4.17.0 --reset-then-reuse-values--reset-then-reuse-values, or your values file with -f. Not --reuse-values. That replays 4.16.0’s values and drops every default this chart added, and the render dies on the first one it needs (nil pointer evaluating interface {}.enabled, from secretKey). Nothing changes on the cluster when it does, so just run it again the right way.
What happens on the first start
Section titled “What happens on the first start”We checked this by running the released 4.16.0 image on a database, then 4.17.0 on the same one.
- Migrations 0009 to 0015 run:
0009_helpers,0010_logbook,0011_scripts,0012_tag_device_class,0013_recipes,0014_blueprintsand0015_unifi. All additive. On an empty plant they took a fifth of a second. On yours they take as long as your database does.pg_dumpfirst, like always. - Your device passwords get encrypted. The chart makes
<release>-secret-keyon this upgrade (it didn’t exist in 4.16.0), and the Gateway seals every stored credential at that start and says so in its log:encrypted stored device credentials devices=N. The devices keep connecting, and nobody has to type anything in again. - From then on, that key is part of your data. Start the Gateway without it and every device with a sealed password stays down, retrying, with an error that names the key it needs. A backup taken after the upgrade restores only where that key is mounted, and
--restore-fromrefuses it by name anywhere else. If the key is gone for good,--restore-keep-unreadable-credentialsbrings everything else back and leaves those devices down until somebody types the passwords in again. A backup from 4.16.0 restores anywhere, and its credentials are sealed as soon as the restore finishes. - Going back to 4.16.0 on the same database starts (we tried it), but 4.16.0 knows nothing about sealing. It hands each device the sealed text as its password, so anything that checks a password refuses it. If you have to go back, restore the
pg_dumpyou took before the upgrade. That’s the whole reason you took it. - Rendering the chart instead of installing it? The key is generated by a
lookupthat only a realhelm installorhelm upgradecan do. Anything that runshelm templateand applies the result gets a new key on every render, and the next sync locks every sealed password out. Make your own key withhotloop gen-secret-key, put it in a Secret you manage, and setsecretKey.existingSecret. Helm, Fleet and the EmberNET App Store all install for real and don’t need this.
So the new job after this upgrade: back up <release>-secret-key the way you back up the rest of the cluster’s secrets, and keep it somewhere other than next to your backups. A key sitting beside the backups it protects protects nothing. Install has the rest of what the key covers.
The Edge Relay
Section titled “The Edge Relay”The relay is 4.17.0 too, released from the same tag. Upgrade its chart with --version 4.17.0 and the same flag rule: --reset-then-reuse-values or your values file. --reuse-values on the relay doesn’t even fail. It drops the chart’s new pushedDevicesFile default, renders the variable empty, and a push goes right back to dying on the read-only mount. Then know four things.
- A pushed device file lives at
/data/pushed-devices.json(HOTLOOP_EDGE_RELAY_PUSHED_FILE, the chart’spushedDevicesFile), on the relay’s volume. It wins over the chart’sdevicesfor as long as it exists, through restarts, reschedules and ahelm upgradethat changesdevices. The relay logs which file it booted from. To go back to the local file,DELETE /api/fleet/{id}/configon the Gateway (fleet.config) while the relay is connected, then restart it. - An old push can come back. A push you made to a 4.16.0 relay never landed, but the broker still holds it, retained, and a 4.17.0 relay picks it up at its first boot and runs it. If it’s stale, clear it with the same
DELETEright after the upgrade and restart the relay. - Quadlet: take the new
edge-relay.container. It pins 4.17.0 and setsHOTLOOP_EDGE_RELAY_PUSHED_FILE. Without that line a push still fails on the read-only root. The Edge Relay install page has the unit in full. persistence.enabled=falseboots now, on an emptyDir capped atpersistence.size. Know what it costs before you pick it: every time the pod is replaced you lose the forward queue, any pushed device file, and the OPC UA client certificate. Anything you care about gets a PVC, which is the default.
What a script against the API will notice
Section titled “What a script against the API will notice”- A time that hasn’t happened is left out. Not
null, left out. 4.16.0 sent0001-01-01T00:00:00ZforlastRunon a rule or script that never ran,lastPollandlastGoodon a device never polled,clearedAt,ackedAtandshelvedUntilon a standing alarm,raisedAton one never raised, the same fields on the live feed’salarmevent, and a few more. That’s how a rule that had never run said “Last run 739887d ago”. If your code tests for year one, test for the field being there instead. - Operators get new permissions the moment you upgrade:
helpers.write,scripts.run,recipes.applyandblueprints.apply. So your operators can set helpers, run scripts, apply recipes and make rules from blueprints straight away. Writing scripts, recipes and blueprints (scripts.write,recipes.write,blueprints.write) is admin. Everybody reads them. A recipe can also name its ownrecipes.apply.<name>. - New routes:
/api/helpers,/api/logbook,/api/scripts,/api/recipes,/api/blueprints,POST /api/automations/{id}/blueprint-upgrade, the UniFi ones (/api/devices/{id}/networkand/api/unifi/{id}/...), andDELETE /api/fleet/{id}/config. - Tags carry
deviceClass, and it rides through to their entities. - Rules carry
blueprintId,blueprintVersionandblueprintInputs. A rule written by hand has none. Saving a rule still replaces it whole, and a body can’t move a rule to a different blueprint or version.
Upgrading from 4.3.1
Section titled “Upgrading from 4.3.1”4.3.1 was published in August, and nothing after it was until 4.16.0. Everything from 4.4.0 to 4.15.3 was merged and never published, so leaving 4.3.1 you take all of it in one jump, under a new name, a new image, a new chart and a new license. This section takes you to 4.16.0, the release this procedure was proven against. From there, Upgrading from 4.16.0 above is the plain upgrade to 4.17.0. Read both before you touch a running plant. They’re short, and they’re shorter than explaining to the plant manager why the historian is empty.
New name on everything
Section titled “New name on everything”| What | 4.3.1 | 4.16.0 and later |
|---|---|---|
| Image | ghcr.io/embernet-ai/industrial-iot |
ghcr.io/hotloop-io/hotloop |
| Chart repository | the old address, now a 404 | https://hotloop.io/hotloop |
| Chart | industrial-iot |
hotloop |
| Environment variables | IIOT_* |
HOTLOOP_* (since 4.9.0) |
It is not a plain helm upgrade
Section titled “It is not a plain helm upgrade”The chart’s name is baked into the Deployment’s and the StatefulSet’s selectors, and Kubernetes flat out refuses to change a selector in place. Point the old release at the new chart and it just fails. Keep the release name and the namespace, and:
-
pg_dumpthe database. Belt and braces. You’ll be glad you did. -
Delete the old Deployment (
<release>) and StatefulSet (<release>-postgresql). The database volume is its own claim, and it stays put. -
Add the repository and upgrade with the values file you ran 4.3.1 with:
Terminal window helm repo add hotloop https://hotloop.io/hotloophelm repo updatehelm upgrade <release> hotloop/hotloop --version 4.16.0 \-n <namespace> -f your-4.3.1-values.yamlNothing in that file was removed from the chart. Rename any
IIOT_inextraEnvtoHOTLOOP_, and drop animage.repositoryoverride if you had one. Don’t use--reuse-values. It skips every default this chart added since 4.3.1, which is more than half its keys, and you’d be running a chart that’s mostly holes.
The names, the Secret’s database password and the volume claim render identical to 4.3.1’s, so the new database pod mounts the old data. The Gateway migrates the schema forward on start, and updates TimescaleDB from 2.29.1 to 2.30.1 while it’s at it.
We proved that by rendering both charts side by side, not by running it on a live cluster. That gap is exactly why step 1 exists. If the upgrade surprises you, you have a dump from five minutes ago, and we want to hear about it at support@hotloop.io.
What behaves differently
Section titled “What behaves differently”- There’s a login. Before 4.4.0 there wasn’t one at all. The chart creates an
adminaccount on first install; Install has the command that reads the password back. - Writes default to off, agents included, since 4.4.0. If your 4.3.1 values file never set
safety.allowWrites, every write your operators and rules used to make is refused after the upgrade until you set it. That’s deliberate, and it’s your cue to review what’s armed before you turn it back on. - A timezone that won’t load stops startup.
HOTLOOP_TIMEZONEorTZset to something bogus used to run on UTC without a word, which quietly shifts every report, schedule and shift change by hours. - Device credentials read back as
[redacted]to every role, admin included (4.15.1). Saving a device with the placeholder keeps the stored password. - The poll settings are
HOTLOOP_POLL_INTERVAL,HOTLOOP_MIN_POLL_INTERVALandHOTLOOP_POLL_TIMEOUT. The old Settings screen told you to setHOTLOOP_POLL_INTERVAL_MS,HOTLOOP_MIN_POLL_INTERVAL_MSandHOTLOOP_TIMEOUT_MS, which nothing ever read. If you set those, they did nothing. Rename them. - OPC UA tags you imported keep the type they were imported with. Discover tags used to call every node a float. It reads each node’s real type now, but tags imported before keep
float, and HotLoop’s own OPC UA server publishes a tag as its declared type, so a text tag reaches Ignition asBad_TypeMismatch. Go declare your text tagsstring. - Edge publish has a values block now. If you set it through
extraEnv, move it toedgePublish:and drop the variables, or they’re set twice. - The OPC UA server says
HotLoopas its manufacturer, where it used to say Fireball Industries. If a client or an inventory script matches on that string, change it before it stops matching. - The license. 4.3.1 was Apache-2.0. From 4.16.0 on it’s the HotLoop Community License v1.0: free for individual, home, hobbyist, nonprofit and education use, and business use is free through EmberNET. See Licensing.
API changes from 4.3.1
Section titled “API changes from 4.3.1”If you have a script or an integration against the API, this is the list. Everything else is additions. The 4.16.0 to 4.17.0 changes above come on top.
- Empty lists are
[], nevernull, in every response and every stream event, at any depth.GET /api/runs, automations, a device’s tags, the write audit, users, alarm events, sparklines and an empty history window all used to answernull. Anullyou still see is never a list; it’s a documented nullable field, like an entity’sequipment_id. /api/historybuckets with no good reading arenull, and draw as gaps, and bad readings stay out of trend averages. A trend and a report over the same hour now say the same thing.- An error the API can’t place is a 500 with a reference into the server log, not a 400. A device that errors on a write or a browse is a 502. The write gate has every status a write can get.
alarm.turn_on, and any other service an alarm entity doesn’t take, is a 400 (“unknown service”). It used to answer success and do nothing, which is the worst possible answer.- Unshelving an alarm that isn’t shelved is a 409. It used to set the alarm to normal whatever it was doing, which could wipe a cleared, unacknowledged alarm off the list, the exact alarm ISA-18.2 exists to keep on it.
- Tag JSON always carries
minValue,maxValueandhasRange, andhasRangealone says whether there’s a range. A range starting at 0 used to look like no minimum. - New:
?dry_run=trueonPOST /api/services/{domain}/{service}. It runs the gate’s checks, writes nothing, audits nothing, and answers with the refusal the real call would give.dry_runtakestrueorfalseand nothing else;yesis a 400, because a guess the wrong way turns a question into a command. GET /api/settingsnames the variable a value actually came from inenvVar, fallbacks included, and a variable set to an empty string counts as unset.
The release notes have the rest of what changed, version by version back to 4.4.0. Once you’re up on 4.16.0 and the historian is still full, go do Upgrading from 4.16.0.