Skip to content

Upgrading HotLoop Gateway to 4.17.0

Two ways onto 4.17.0, and they are nothing alike. From 4.16.0 it’s a plain helm upgrade with one flag that bites, and afterwards a new Secret is as much a part of your plant’s data as the database is. From 4.3.1 it’s a renamed chart that a plain helm upgrade can’t touch, and you go through 4.16.0 on the way. Find the one that’s yours and read all of it before you touch a running plant.

Why bother: 4.16.0 has bugs a running plant can trip over, and 4.17.0 fixes them. The command prompt can swallow a Set, so an operator watches the prompt close, figures the setpoint went out, and nothing was written. Saving a rule in the editor, even to rename it, quietly resets its run budget to five minutes. A device file pushed to an Edge Relay never lands, while the Gateway says config pushed. And an EmberNET App Store deploy loses its tenant. On top of the fixes you get recipes, scripts, blueprints, helpers, the logbook, the new automation steps, read-only UniFi, and MCP going from 16 tools to 25.

It’s a plain helm upgrade. We rendered both charts at both versions: every selector and every volume claim is identical, and the only new object is the key’s Secret.

Terminal window
helm repo update hotloop
helm upgrade <release> hotloop/hotloop --version 4.17.0 --reset-then-reuse-values

--reset-then-reuse-values, or your values file with -f. Not --reuse-values. That replays 4.16.0’s values and drops every default this chart added, and the render dies on the first one it needs (nil pointer evaluating interface {}.enabled, from secretKey). Nothing changes on the cluster when it does, so just run it again the right way.

We checked this by running the released 4.16.0 image on a database, then 4.17.0 on the same one.

  • Migrations 0009 to 0015 run: 0009_helpers, 0010_logbook, 0011_scripts, 0012_tag_device_class, 0013_recipes, 0014_blueprints and 0015_unifi. All additive. On an empty plant they took a fifth of a second. On yours they take as long as your database does. pg_dump first, like always.
  • Your device passwords get encrypted. The chart makes <release>-secret-key on this upgrade (it didn’t exist in 4.16.0), and the Gateway seals every stored credential at that start and says so in its log: encrypted stored device credentials devices=N. The devices keep connecting, and nobody has to type anything in again.
  • From then on, that key is part of your data. Start the Gateway without it and every device with a sealed password stays down, retrying, with an error that names the key it needs. A backup taken after the upgrade restores only where that key is mounted, and --restore-from refuses it by name anywhere else. If the key is gone for good, --restore-keep-unreadable-credentials brings everything else back and leaves those devices down until somebody types the passwords in again. A backup from 4.16.0 restores anywhere, and its credentials are sealed as soon as the restore finishes.
  • Going back to 4.16.0 on the same database starts (we tried it), but 4.16.0 knows nothing about sealing. It hands each device the sealed text as its password, so anything that checks a password refuses it. If you have to go back, restore the pg_dump you took before the upgrade. That’s the whole reason you took it.
  • Rendering the chart instead of installing it? The key is generated by a lookup that only a real helm install or helm upgrade can do. Anything that runs helm template and applies the result gets a new key on every render, and the next sync locks every sealed password out. Make your own key with hotloop gen-secret-key, put it in a Secret you manage, and set secretKey.existingSecret. Helm, Fleet and the EmberNET App Store all install for real and don’t need this.

So the new job after this upgrade: back up <release>-secret-key the way you back up the rest of the cluster’s secrets, and keep it somewhere other than next to your backups. A key sitting beside the backups it protects protects nothing. Install has the rest of what the key covers.

The relay is 4.17.0 too, released from the same tag. Upgrade its chart with --version 4.17.0 and the same flag rule: --reset-then-reuse-values or your values file. --reuse-values on the relay doesn’t even fail. It drops the chart’s new pushedDevicesFile default, renders the variable empty, and a push goes right back to dying on the read-only mount. Then know four things.

  • A pushed device file lives at /data/pushed-devices.json (HOTLOOP_EDGE_RELAY_PUSHED_FILE, the chart’s pushedDevicesFile), on the relay’s volume. It wins over the chart’s devices for as long as it exists, through restarts, reschedules and a helm upgrade that changes devices. The relay logs which file it booted from. To go back to the local file, DELETE /api/fleet/{id}/config on the Gateway (fleet.config) while the relay is connected, then restart it.
  • An old push can come back. A push you made to a 4.16.0 relay never landed, but the broker still holds it, retained, and a 4.17.0 relay picks it up at its first boot and runs it. If it’s stale, clear it with the same DELETE right after the upgrade and restart the relay.
  • Quadlet: take the new edge-relay.container. It pins 4.17.0 and sets HOTLOOP_EDGE_RELAY_PUSHED_FILE. Without that line a push still fails on the read-only root. The Edge Relay install page has the unit in full.
  • persistence.enabled=false boots now, on an emptyDir capped at persistence.size. Know what it costs before you pick it: every time the pod is replaced you lose the forward queue, any pushed device file, and the OPC UA client certificate. Anything you care about gets a PVC, which is the default.
  • A time that hasn’t happened is left out. Not null, left out. 4.16.0 sent 0001-01-01T00:00:00Z for lastRun on a rule or script that never ran, lastPoll and lastGood on a device never polled, clearedAt, ackedAt and shelvedUntil on a standing alarm, raisedAt on one never raised, the same fields on the live feed’s alarm event, and a few more. That’s how a rule that had never run said “Last run 739887d ago”. If your code tests for year one, test for the field being there instead.
  • Operators get new permissions the moment you upgrade: helpers.write, scripts.run, recipes.apply and blueprints.apply. So your operators can set helpers, run scripts, apply recipes and make rules from blueprints straight away. Writing scripts, recipes and blueprints (scripts.write, recipes.write, blueprints.write) is admin. Everybody reads them. A recipe can also name its own recipes.apply.<name>.
  • New routes: /api/helpers, /api/logbook, /api/scripts, /api/recipes, /api/blueprints, POST /api/automations/{id}/blueprint-upgrade, the UniFi ones (/api/devices/{id}/network and /api/unifi/{id}/...), and DELETE /api/fleet/{id}/config.
  • Tags carry deviceClass, and it rides through to their entities.
  • Rules carry blueprintId, blueprintVersion and blueprintInputs. A rule written by hand has none. Saving a rule still replaces it whole, and a body can’t move a rule to a different blueprint or version.

4.3.1 was published in August, and nothing after it was until 4.16.0. Everything from 4.4.0 to 4.15.3 was merged and never published, so leaving 4.3.1 you take all of it in one jump, under a new name, a new image, a new chart and a new license. This section takes you to 4.16.0, the release this procedure was proven against. From there, Upgrading from 4.16.0 above is the plain upgrade to 4.17.0. Read both before you touch a running plant. They’re short, and they’re shorter than explaining to the plant manager why the historian is empty.

What 4.3.1 4.16.0 and later
Image ghcr.io/embernet-ai/industrial-iot ghcr.io/hotloop-io/hotloop
Chart repository the old address, now a 404 https://hotloop.io/hotloop
Chart industrial-iot hotloop
Environment variables IIOT_* HOTLOOP_* (since 4.9.0)

The chart’s name is baked into the Deployment’s and the StatefulSet’s selectors, and Kubernetes flat out refuses to change a selector in place. Point the old release at the new chart and it just fails. Keep the release name and the namespace, and:

  1. pg_dump the database. Belt and braces. You’ll be glad you did.

  2. Delete the old Deployment (<release>) and StatefulSet (<release>-postgresql). The database volume is its own claim, and it stays put.

  3. Add the repository and upgrade with the values file you ran 4.3.1 with:

    Terminal window
    helm repo add hotloop https://hotloop.io/hotloop
    helm repo update
    helm upgrade <release> hotloop/hotloop --version 4.16.0 \
    -n <namespace> -f your-4.3.1-values.yaml

    Nothing in that file was removed from the chart. Rename any IIOT_ in extraEnv to HOTLOOP_, and drop an image.repository override if you had one. Don’t use --reuse-values. It skips every default this chart added since 4.3.1, which is more than half its keys, and you’d be running a chart that’s mostly holes.

The names, the Secret’s database password and the volume claim render identical to 4.3.1’s, so the new database pod mounts the old data. The Gateway migrates the schema forward on start, and updates TimescaleDB from 2.29.1 to 2.30.1 while it’s at it.

We proved that by rendering both charts side by side, not by running it on a live cluster. That gap is exactly why step 1 exists. If the upgrade surprises you, you have a dump from five minutes ago, and we want to hear about it at support@hotloop.io.

  • There’s a login. Before 4.4.0 there wasn’t one at all. The chart creates an admin account on first install; Install has the command that reads the password back.
  • Writes default to off, agents included, since 4.4.0. If your 4.3.1 values file never set safety.allowWrites, every write your operators and rules used to make is refused after the upgrade until you set it. That’s deliberate, and it’s your cue to review what’s armed before you turn it back on.
  • A timezone that won’t load stops startup. HOTLOOP_TIMEZONE or TZ set to something bogus used to run on UTC without a word, which quietly shifts every report, schedule and shift change by hours.
  • Device credentials read back as [redacted] to every role, admin included (4.15.1). Saving a device with the placeholder keeps the stored password.
  • The poll settings are HOTLOOP_POLL_INTERVAL, HOTLOOP_MIN_POLL_INTERVAL and HOTLOOP_POLL_TIMEOUT. The old Settings screen told you to set HOTLOOP_POLL_INTERVAL_MS, HOTLOOP_MIN_POLL_INTERVAL_MS and HOTLOOP_TIMEOUT_MS, which nothing ever read. If you set those, they did nothing. Rename them.
  • OPC UA tags you imported keep the type they were imported with. Discover tags used to call every node a float. It reads each node’s real type now, but tags imported before keep float, and HotLoop’s own OPC UA server publishes a tag as its declared type, so a text tag reaches Ignition as Bad_TypeMismatch. Go declare your text tags string.
  • Edge publish has a values block now. If you set it through extraEnv, move it to edgePublish: and drop the variables, or they’re set twice.
  • The OPC UA server says HotLoop as its manufacturer, where it used to say Fireball Industries. If a client or an inventory script matches on that string, change it before it stops matching.
  • The license. 4.3.1 was Apache-2.0. From 4.16.0 on it’s the HotLoop Community License v1.0: free for individual, home, hobbyist, nonprofit and education use, and business use is free through EmberNET. See Licensing.

If you have a script or an integration against the API, this is the list. Everything else is additions. The 4.16.0 to 4.17.0 changes above come on top.

  • Empty lists are [], never null, in every response and every stream event, at any depth. GET /api/runs, automations, a device’s tags, the write audit, users, alarm events, sparklines and an empty history window all used to answer null. A null you still see is never a list; it’s a documented nullable field, like an entity’s equipment_id.
  • /api/history buckets with no good reading are null, and draw as gaps, and bad readings stay out of trend averages. A trend and a report over the same hour now say the same thing.
  • An error the API can’t place is a 500 with a reference into the server log, not a 400. A device that errors on a write or a browse is a 502. The write gate has every status a write can get.
  • alarm.turn_on, and any other service an alarm entity doesn’t take, is a 400 (“unknown service”). It used to answer success and do nothing, which is the worst possible answer.
  • Unshelving an alarm that isn’t shelved is a 409. It used to set the alarm to normal whatever it was doing, which could wipe a cleared, unacknowledged alarm off the list, the exact alarm ISA-18.2 exists to keep on it.
  • Tag JSON always carries minValue, maxValue and hasRange, and hasRange alone says whether there’s a range. A range starting at 0 used to look like no minimum.
  • New: ?dry_run=true on POST /api/services/{domain}/{service}. It runs the gate’s checks, writes nothing, audits nothing, and answers with the refusal the real call would give. dry_run takes true or false and nothing else; yes is a 400, because a guess the wrong way turns a question into a command.
  • GET /api/settings names the variable a value actually came from in envVar, fallbacks included, and a variable set to an empty string counts as unset.

The release notes have the rest of what changed, version by version back to 4.4.0. Once you’re up on 4.16.0 and the historian is still full, go do Upgrading from 4.16.0.