Commit Graph
18 Commits
Author SHA1 Message Date
tfaour 1d39e439ff Drop the last legacy widget-system and shared-token auth scaffolding
Firmware build check / build-check (push) Successful in 5m37s
Build and release firmware / build-and-release (push) Successful in 5m36s
Build and push server image / test (push) Successful in 1m37s
Build and push server image / build-and-push (push) Successful in 4m18s
Build and push server image / deploy (push) Failing after 1m20s
Server: migration 41 drops the pre-widget-system Frame columns
(mode/album_id/current_asset_id/queue/calendar_*/whiteboard_*, etc)
docs/widgets.md flagged as the deliberately-deferred Phase 6 cleanup,
with a raw-SQL backfill safety net for any frame that still somehow
lacks a Widget. Also drops legacy_token_enabled and the shared
MANAGEMENT_TOKEN fallback it gated in require_device/require_browser --
the per-frame manage_token/device_token flow (and the /m/ page) fully
supersede it now; MANAGEMENT_TOKEN's only remaining role is the
optional pre-setup claim gate. Confirmed with the maintainer that the
deployed frame is already off the shared token before removing the
server-side fallback.

Firmware: the captive portal's "Access Token" field and its NVS/
build_url plumbing only ever mattered for pointing new firmware at an
old pre-multi-frame server -- gone along with the server-side fallback
it fed. Version bump to publish the change.
2026-08-04 18:33:29 +00:00
tfaour fcf3aec4c0 Move button actions to per-widget config, add hold-for-global-action
Build and push server image / test (push) Successful in 36s
Firmware build check / build-check (push) Successful in 2m4s
Build and push server image / build-and-push (push) Successful in 3m12s
Build and push server image / deploy (push) Successful in 58s
Next/back button assignment moves from a frame-level "Button
assignments" card into each widget's own gear-icon dialog, prefilled
with a sane default at creation (photos/calendar -> advance/back,
whiteboard/weather -> check_now, others -> none). At most one binding
per (widget, button) now -- cross-widget execution order never
mattered since each widget's action only touches its own state.

New firmware capability: holding NEXT or BACK past a configurable
duration (min 3s, server-side default) triggers a frame-wide action
instead of the per-widget short-press one -- cycling saved layouts,
refreshing all widgets, or freezing/unfreezing every photo widget (see
app/global_actions.py). Firmware next/back checks gain the same
hold-duration polling the combo button already had; the threshold
comes from the previous wake's /frame/config fetch (persisted in NVS),
since this wake's button decision happens before that request.

Not done here: firmware/version.txt is intentionally left unbumped --
this hasn't been built or hardware-tested (no ESP-IDF toolchain in this
environment), so no firmware release build should be triggered yet.
2026-07-27 22:09:33 +00:00
tfaour 462b558bef Fix captive portal redirect: visible countdown, keep AP up until it finishes
Previously the softAP was torn down (esp_restart) only 1s after sending
the success page, while the page's own redirect timer waited 7s -- so
the AP (and the phone's captive-portal session with it) was gone long
before the redirect could fire. Now the page shows a live 10s countdown
before redirecting, and the device holds the AP up for 11s so the
countdown always completes. Also added a "Redirect now" button for a
phone that's already reconnected to normal WiFi.
2026-07-22 09:45:59 -04:00
tfaour 683e3881b1 Redesign phase C: claim flow, limited manage page, device protocol
The frame-claiming pipeline, end to end. Firmware: every request now
carries ?id=<12-hex STA MAC> via build_url (mirrored in build_ota_url),
and the captive portal's success page became a redirect that hands the
user's browser to <server>/claim?device_id=... after ~7s -- enough time
for the phone to drop the provisioning AP while the device reboots.
The server pushes a per-frame device token through /frame/config during
a one-time handshake; the firmware persists it to NVS (a dedicated
single-key write that deliberately doesn't reset the connected-once
flag or WiFi cache) and prefers it over the provisioned shared token
from the next request on. Config response buffer grows 256->512. Both
board variants compile clean; new firmware also works against an old
server (which ignores ?id=) and old firmware against this server (the
phase A legacy mapping), so either deploy order survives.

Server: /claim lands the captive-portal redirect -- claim-gated signup
(a valid unclaimed/unregistered device id IS the enrollment invitation),
pending claims for the user-beats-the-frame race (auto-attached at
self-registration, 24h expiry), and a waiting page that refreshes until
the frame checks in. Unclaimed/unconfigured frames get a rendered
instruction placeholder with a QR from /frame/image (200, never an
error loop) -- new qrcode dep, placeholder shares the exact
quantize/pack path photos use.

The on-frame manage QR now resolves to a limited no-login page: scans
of / carrying device credentials (new ?id&token or the legacy shared
token) 303 to /m/<manage_token>, which allows exactly view queue,
show-next, advance, back, and scoped thumbnails -- no settings, no
removal, no other frames. Full control means logging in.

One real protocol hole found by simulating full wake cycles: after
self-registration the device could never authenticate again (the wake
cycle fetches the image BEFORE /frame/config delivers its token).
require_device now treats the id itself as the credential until the
first authenticated request flips device_token_ack -- the same trust
level as open registration, closing permanently once the handshake
completes.
2026-07-21 23:44:22 -04:00
tfaour 4f4b2844e6 Firmware: WiFi fast-connect cache (skip scan + DHCP on the next wake)
After a successful home-WiFi connection, caches BSSID/channel and
IP/netmask/gateway/DNS in NVS. The next wake's first connect attempt
uses the cached BSSID/channel (skips the all-channel scan) and applies
the cached IP directly once the link comes up (skips DHCP) -- a couple
fewer seconds of radio-on time per wake, free every wake since nothing
about the network actually needs renegotiating most of the time.

Falls back to a normal scan+DHCP attempt, and clears the cache, if: the
fast attempt itself fails, or it "succeeds" at the WiFi layer but the
full fetch cycle then fails anyway (a stale cached IP/DNS/gateway that
associates but can't actually reach the server). Also cleared on
(re)provisioning and factory reset, since a new network shouldn't try
to reuse the old one's cache.

The static-IP path needed care to get right without touching untested
territory: esp_netif_set_ip_info() only posts IP_EVENT_STA_GOT_IP (what
the existing connect-wait logic blocks on) once the netif is already
up, which the internal netif-glue's own WIFI_EVENT_STA_CONNECTED
handler guarantees by running first (registered earlier, in
esp_netif_create_default_wifi_sta()) -- confirmed against ESP-IDF's own
static_ip example and esp_netif_handlers.c source rather than assumed.
Falling back after a failed fast attempt also needed an explicit
esp_netif_dhcpc_start() first: esp_netif_dhcpc_stop() leaves the netif's
DHCP status STOPPED rather than resetting to INIT, and left alone the
glue would silently re-post the stale cached IP on the next connect
instead of actually running DHCP (esp_netif_action_connected).

Version bumped to 1.1.0 (real feature, not just a fix); build-verified
clean on both board configs (devkit 8MB, XIAO 4MB), no new warnings.
2026-07-20 23:46:57 -04:00
tfaour a3ab6c5f13 Firmware: OTA client, dual-board build (devkit/XIAO), version reporting, XIAO fixes
Build and push server image / build-and-push (push) Successful in 36s
- version.txt + esp_app_desc_t version reporting (X-Frame-Version header);
  new ota_update.c checks the server's advertised version against the
  running one and streams+applies an update via esp_https_ota, gated by
  bootloader rollback (marks the image valid only after a full successful
  cycle, so a bad update can't brick a wall-mounted frame).
- Dual-OTA partition tables: partitions.csv (8MB dev board, 2MB slots) and
  new partitions_xiao.csv (4MB XIAO, 1.875MB slots -- the dev board's
  table doesn't fit the XIAO's flash). New build_for_board.sh gives each
  board its own build dir + generated sdkconfig via SDKCONFIG_DEFAULTS
  layering, so switching boards never clobbers the other's config.
- fetch_photo_info()/fetch_face_labels() were using the short
  reachability-check timeout even though the manage-menu path can be the
  first (cold, TLS-handshake-paying) request of a wake cycle -- switched
  to the longer fetch timeout to stop spurious ESP_ERR_HTTP_CONNECT
  failures.
- XIAO: the RF switch that selects onboard vs. external antenna
  (GPIO3/14) isn't initialized by plain ESP-IDF the way Seeed's Arduino
  package does it, leaving WiFi unable to reliably reach the antenna at
  all -- new board_antenna.c powers the switch and selects the onboard
  antenna, gated behind FRAME_XIAO_ANTENNA_INIT (on by default in
  sdkconfig.xiao). Also remaps the EPD DC/RST/BUSY pins, since the dev
  board's defaults (GPIO9/10/11) aren't physically exposed on the XIAO.
2026-07-20 22:24:15 -04:00
tfaour 6c7468a36e Add HTTPS support and a management-token gate for the web UI
Build and push server image / build-and-push (push) Successful in 31s
ESP32 side can now reach the tools server over HTTPS: the Tools Server
field accepts an https:// address for a TLS-terminating reverse proxy
in front of the server (which still only ever speaks plain HTTP
itself), trusting Cloudflare's Origin CA root (embedded at build time)
since that's the common way to get a real cert on a private origin.
Every URL the device builds -- image fetch, config check, manage-menu
data, the QR codes' own links -- goes through one build_url() helper
that picks the scheme from what's configured.

Also adds an optional MANAGEMENT_TOKEN (docker-compose.yml) that gates
the web UI (/, /api/*) behind a shared secret -- unset by default, so
existing trusted-LAN deployments are unaffected. The same token is
entered once during the ESP32's captive-portal setup and gets baked
into the manage-menu's QR code (?token=...), so scanning it just works;
visiting the page without a valid token shows a plain entry prompt
instead of the config UI, and a valid query-param hit sets a cookie so
the page's own fetch()/<img> calls stay authorized for the rest of the
visit. Device-facing /frame/* endpoints are unaffected -- a separate,
already-documented trust boundary.
2026-07-19 09:42:38 -04:00
tfaour c4cd9b73e8 Invalidate the tracked display CRC when a non-photo screen is drawn
The QR onboarding and "CONNECTING..." status screens write to the panel
through a separate path that never touched the last-displayed-photo CRC
added in the previous commit. That left it stale relative to what's
actually on screen after either one draws -- most visibly after a
factory reset: reprovisioning and reconnecting could fetch a photo whose
CRC happened to match the one from before the reset, skip the refresh,
and leave the QR code frozen on screen indefinitely. Both screens now
invalidate the tracked CRC right after drawing, so the next photo fetch
is always guaranteed to actually refresh.
2026-07-18 23:59:48 -04:00
tfaour b9649c35ec Skip redundant panel refreshes and fetch the image before the config check
The panel driver now splits writing a frame into its SPI buffer
(epd_write_frame(), which also computes a CRC32 as it streams) from
actually triggering the physical refresh (epd_turn_on_display()).
frame_client.c compares the new CRC against the last one that was
actually refreshed (persisted in NVS) and skips the refresh entirely
when they match -- e.g. a reboot redisplaying the same photo before the
server's refresh interval elapsed no longer causes a visible flash for
no visual change.

Also reorders the per-wake fetch cycle: the image fetch (15s timeout)
now goes before the config check (3s timeout), instead of after. The
config check's tighter timeout was intermittently tripping on
connection-setup latency that's common on the first request after
waking from a long deep sleep (e.g. stale ARP); putting the more
tolerant request first absorbs that latency, and the config check then
rides the connection it already warmed up.
2026-07-18 23:51:01 -04:00
tfaour d395cf3bb9 Add two physical buttons: factory-reset and next-photo
Build and push server image / build-and-push (push) Successful in 35s
Factory-reset (GPIO3, hold 10s): clears stored WiFi/server config and
restarts into provisioning -- the deliberate, USB-free replacement for
the earlier reverted RST-based auto-reprovisioning idea.

Next-photo (GPIO2, tap): wakes the device and forces the server to
advance immediately via a new POST /frame/advance, instead of waiting
for the refresh interval. Both buttons arm themselves as deep-sleep GPIO
wakeup sources so a press is noticed promptly even while asleep.

Also makes GET /frame/image side-effect-free: it now only advances once
refresh_interval_s has elapsed since the current photo was set (tracked
server-side), so a device reboot for any reason just redisplays the
current photo instead of silently skipping ahead. The server maintains a
small reorderable upcoming-photos queue, viewable and rearrangeable from
the web UI.
2026-07-18 23:28:36 -04:00
tfaour 012c6dda8e Revert "Add two ways back into provisioning: RST press and repeated server failure"
This reverts commit d32236d832.
2026-07-18 16:24:38 -04:00
tfaour d32236d832 Add two ways back into provisioning: RST press and repeated server failure
RST/power-on trigger: checks esp_reset_reason() at the very top of boot.
ESP32-C6 can't electrically distinguish the RST/EN button from a genuine
power-on (both report ESP_RST_POWERON -- confirmed against ESP-IDF's own
docs, ESP_RST_EXT is explicitly "not applicable"), so POWERON is treated
as "user wants to reconfigure" and routes straight to provisioning. Safe
because the device's only normal restart path is ESP_RST_DEEPSLEEP (its
own scheduled wake), and crash-type resets (brownout/watchdog/panic)
report their own distinct reasons, not POWERON -- so a flaky power supply
or transient crash won't get bounced into provisioning, only an actual
power cycle or RST press will (which plausibly means the frame is being
moved/redeployed anyway).

Auto-fallback: a new NVS-persisted consecutive-failure counter
(frame_config_record_server_failure/reset_server_failures) tracks wakes
where the tools server was unreachable. After
CONFIG_FRAME_REPROVISION_AFTER_FAILURES in a row (default 12, ~1hr at the
retry interval), the device clears its stored WiFi config and
esp_restart()s rather than calling wifi_provisioning_start() directly --
doing that inline would mean initializing the display driver a second
time in the same session (frame_client_run already did once), the same
class of double-init bug hit earlier with WiFi. The next boot's
frame_config_load() naturally reports "not provisioned" and routes
through the existing, already-tested provisioning path with a single
fresh epd_init(). Solves the "I moved the server to a new address" case
without needing USB access.

Both frame_config_save() (fresh provisioning) and any successful server
contact reset the failure counter.
2026-07-18 16:10:03 -04:00
tfaour 878909b302 Fix status screen logic: first-connection screen, false-negative skip
Two bugs found testing against a real (partially-configured) server:

- The reachability check hit HEAD / with -- our server only registers
  GET on that route, so it always got a 405. Harmless for the check
  itself (any completed HTTP response counts as "reachable"), but noisy
  and semantically wrong. Points at GET /health instead, which exists
  for exactly this.

- fetch_and_display() failing before any pixel data was sent (e.g. a
  400/404 on /frame/image) was treated the same as a mid-stream failure,
  which skips drawing a status screen to avoid compounding flashing on
  top of an already-refreshed panel. But a pre-stream failure never
  touches the panel at all, so skipping the status screen there just
  left the old provisioning QR code on screen with no indication
  anything had gone wrong. fetch_and_display() now reports whether
  streaming ever started so the caller can tell the two cases apart.

- The status screen now always shows on the very first successful
  connection after (re)provisioning, regardless of outcome, via a new
  "connected_once" NVS flag that frame_config_save() resets on every
  fresh provisioning event. Later wakes skip it on success (straight to
  the photo) but still show it on any failure, matching the intent from
  the original status-screen feature.
2026-07-18 15:06:50 -04:00
tfaour ecc42f23c0 Redesign setup screen: title + two-step layout with a config QR
Adds a "2. CONFIGURATION" step alongside WiFi setup: a second QR code
linking straight to the captive portal page (http://<ap-ip>/), for anyone
who's joined the AP but wants a one-scan shortcut to the config form
instead of relying on the captive-portal popup. The AP netif is now
created (but not started) before the display renders, since the AP's IP
is fixed at netif creation and needed for this QR before the network is
actually up.

Pulls the pixel/text drawing primitives (previously private to
qr_onboarding.c) out into epd_draw.c/.h so the upcoming status screen can
reuse them instead of duplicating.
2026-07-18 14:07:56 -04:00
tfaour 1dd02da70a Add QR-code WiFi onboarding screen
Vendors two small MIT/BSD-3-Clause libraries rather than hand-rolling
either: Nayuki's qrcodegen (QR matrix generation) and Waveshare's Font24
bitmap table from their e-Paper repo (same repo the epd7in3e driver came
from) for rendering readable text on the panel.

qr_onboarding_show() builds a standard WIFI:T:WPA;S:...;P:...;; payload,
rasterizes the QR module matrix plus the SSID and password as plaintext
underneath (for anyone provisioning from a device that can't scan a QR)
onto a malloc'd frame buffer, and pushes it to the panel via
epd_display_buffer(). The buffer is heap-allocated on demand rather than
statically reserved, since 192KB held permanently in BSS would eat into
the RAM budget the HTTP fetch path (task 6) is specifically trying to keep
free.

Wired into wifi_provisioning_start() before the softAP comes up, so the
join instructions are already on-screen by the time the network is
joinable.
2026-07-18 11:20:43 -04:00
tfaour 85a5238724 Add epd7in3e display driver component
Ports Waveshare's official EPD_7in3e.c register/refresh sequence (the
panel has no public datasheet, so their reference driver is the source of
truth) to an ESP-IDF component using spi_master + gpio instead of the
bcm2835/RPi hardware abstraction the reference targets.

Unlike the reference driver, which toggles CS around every single byte,
this holds CS low for each logical command/data phase and DMAs pixel data
in 4KB chunks -- sending ~192,000 individual one-byte SPI transactions
would make a full refresh impractically slow.

Exposes a streaming API (epd_display_stream, pulling chunks from a
caller-supplied read_fn) so the eventual HTTP fetch path can feed the panel
without holding a full ~192KB frame in RAM, plus an in-memory convenience
wrapper (epd_display_buffer) for cases like the upcoming QR code screen
where buffering the whole frame is fine.

Wired into wifi_provisioning_start() as an init + white-clear for now, to
prove the driver builds/links/runs at the right point in the boot sequence
ahead of the actual QR code content.
2026-07-18 11:05:06 -04:00
tfaour 87c2eccfdf Suffix provisioning AP SSID with device MAC to avoid collisions
Two frames on the same network would otherwise both advertise the same
ESPRESSO SSID during provisioning. Appends the last 3 bytes of the WiFi MAC
(read via esp_read_mac(), available before the WiFi driver starts) so each
device's softAP is uniquely named.
2026-07-18 10:52:56 -04:00
tfaour de449d5973 Finish captive portal: NVS config, POST handler, STA connect, AP identity
Splits provisioning (softAP + captive portal + NVS-backed config) into
wifi_provisioning.c and home-network connection into frame_client.c.
/save_config now parses the form body and persists it to NVS; on boot the
device goes straight to STA mode if a config exists, retrying a few times
before falling back to provisioning if the home network is unreachable.

The provisioning AP is now always named ESPRESSO with a random per-device
password (generated once, persisted in NVS) instead of a fixed Kconfig
value, drawn from a charset that avoids visually ambiguous characters since
it'll be read off the e-ink panel and possibly typed by hand.
2026-07-18 10:48:52 -04:00