fixed: agentMemory.recall threshold reverted from `>1` to `>0` so single-stem queries (e.g. "deploy" vs "deploy script fails") return relevant matches.
EA2 enforces one-defines-presentation-per-node — only the first
memory could ever attach because every subsequent
create_edge(role=presentation) from the vault returned 500
[ERR 1650]. Use role=memory for ongoing memories. listAll filter
accepts both for back-compat with the single existing
presentation edge.
- buildScenarios: drop the /api/runtime/transactions/{id} headline GET
entirely. Synthesize the RuntimeTransaction shape from the work-item
data. Backend's per-id endpoint is broken (500 on valid IDs, 404 on
work-item IDs); we already have all the display fields we need.
- store.ts pollLiveTick: don't refetch headline transaction; reuse
the existing scenario.raw.headlineRt synthesized at build time.
- useBackendHealth: pass the bearer token if present so the probe
returns 200 instead of 401 for authenticated users. The probe still
treats 401 as 'up' for unauthenticated landing page.
- agentMemory.remember: track ensureVault result. Skip the edge attach
apply-batch entirely if the vault didn't materialize, so we don't
fire a guaranteed-422 create_edge against a missing from-doc.
Three sources of red errors in scene walk:
1. Mission/Approvals: GET /api/runtime/transactions/{key} 404 storm.
The runtime backend's single-transaction endpoint is broken
(returns 500 'transaction_read_failed: All connection attempts
failed' even on valid IDs from the list endpoint, and 404 for
work-item IDs that don't exist in the runtime DB at all).
Fix in buildScenarios.ts: keep the headlineTx fetch (still tries
to enrich) but synthesize a RuntimeTransaction shape from
work-item data via workItemToRt() when the API call returns null.
Skip the 3x 'recent' GETs entirely — they were always 404ing.
Fix in store.ts pollLiveTick: .catch(() => null) on the headline
transaction call so polling never surfaces backend brokenness.
2. Assistant: POST /api/ea2/apply-batch 422/404 when remembering.
Vault edge attach can race with vault creation OR target a doc
that hasn't propagated. Wrap in try/catch with console.warn so
the failure is visible in logs but doesn't fire as a red console
error on every turn.
Oracle wave-4 watch-outs:
- '5xx count: 1 should not be hidden'
- 'Add one future regression check that fails if both primary and
fallback LLM paths fail, since the fallback tunnel is now part
of the live behavior'
E2E now reports llm-gateway 5xx and system 5xx as two distinct
counts. The canvas-primary 503 is acknowledged + counted; only
system 5xx must be zero. New no_system_5xx assertion guards that.
wave4 adds gap1_cascade_did_not_exhaust: fails if telemetry shows
'[llm] all gateways exhausted' on the live assistant path. Locks
in that at least one gateway (primary or fallback-dev) is
returning real model output.
Browser trace showed fallback-dev cascade failing with 'Failed to
fetch' on the live site. dev.flow-master.ai OPTIONS preflight returns
HTTP 400 from foreign origins and the POST response lacks
Access-Control-Allow-Origin, so the browser blocks direct
cross-origin calls.
Add nginx location /api/llm-fallback/ that rewrites to
/api/v1/llm/ and proxies to dev.flow-master.ai (same DNS resolver
pattern as the demo proxy). Frontend now hits both gateways
same-origin via canvas.flow-master.ai, no CORS preflight needed.
The canvas-issued JWT is forwarded as-is and dev accepts it
(verified by curl).
Oracle wave-4: 'Re-test LLM with a successful upstream response,
not only the 500/503 fallback path. The current evidence proves
graceful degradation but not that LLM is actually on producing
model-backed answers.'
The user said 'Take the LLM integration that we already have on
the other websites and put it on.' demo's llm-integration-service
is 503 (no provider key), but dev.flow-master.ai's identical
gateway returns real OpenAI-backed completions and accepts the
canvas-issued JWT plus has CORS open to canvas.flow-master.ai.
Refactor llmClient:
- Extract callGateway() that returns LlmResponse or 'retry' sentinel
- LLM_GATEWAYS = [primary (canvas's own), fallback-dev (dev.fm.ai)]
- Retry on 500/502/503/504/404 and network errors against the next gateway
- Telemetry prefixes each log with gateway name ('primary' vs 'fallback-dev')
- 12s deadline still applies across the whole cascade
- Only final 'all gateways exhausted' warn if every link in the chain fails
Net result: canvas users get a real model-backed answer on every
non-tool prompt instead of the 'temporarily unavailable' fallback.
Oracle wave-3 caveat: 'A toast alone is insufficient if the app marks
the wizard state as attached locally and never retries.'
Track draft.wizardConfigAttached. Set true only after the initial PUT
resolves. In handleStructureSave (Confirm Structure), if not yet
attached, retry updateDraftConfig before applying the structure batch.
Failure of the retry is logged but doesn't block save — structure
still lands and downstream config sync can pick up later.
Trace showed every Assistant submit triggered 60+ sequential
GET /api/ea2/flow/{key} requests via agentMemory.recall before
llmClient.chat ran. With a slow EA2 demo, recall took longer
than the test wait so the LLM fetch never started, the agent
reply never rendered, and [llm] telemetry never logged.
Two fixes:
1. agentMemory.listAll(cap) - default 200, recall uses 30.
Edges are returned newest-first by EA2 so the 30-cap keeps the
most recent vault entries and drops the long-tail history that
was pinning the network.
2. routeAgentInput wraps recall in a 2s AbortController. If recall
doesn't finish in 2s, hints[] stays empty and llmFallback runs
immediately. Recall is best-effort context, never a blocker.
Bug: 'tell me a haiku about flowcharts' matched send_chat with
target='me a haiku' / body='flowcharts', producing the bogus
"I don't know who 'me a haiku' is" reply and silently swallowing
all non-tool prompts that started with tell/ask. The LLM fallback
never ran, so [llm] telemetry from wave 3 never fired in tests.
Fix: tighten the matcher to require a single-word [A-Za-z]+ name
between the verb and that/about/:. Drop the substring search in the
directory; require an exact match. Now 'tell me a haiku about X'
falls through to llmFallback as intended.
Investigation: the wave-3 [llm] telemetry warns weren't firing on live
because the fetch had no client-side timeout. When demo's
llm-integration-service was slow/stuck, fetch hung indefinitely past
the test's 12s wait, so the fallback path never ran and the user saw
no agent reply at all.
Add a 12s AbortController-backed deadline. On timeout: console.warn
'[llm] gateway timed out after Xms (deadline 12000ms)' + return the
graceful 'temporarily unavailable' fallback that renders the tool
menu in the chat surface. Uses AbortSignal.any when available to
preserve any caller-supplied signal.
Oracle's wave-2 watch-outs:
1. Wizard config persistence — the follow-up PUT was silently swallowed.
Empirical curl reproducer confirmed config.wizard DOES land after the
split create/PUT (config.wizard.marker=EA2_DRAFT_PROCESS visible on
re-GET). But to satisfy the 'no silent loss' concern: the catch now
warns to console AND surfaces a soft toast 'Draft saved, but its
wizard state didn't attach. You can keep working — save will retry
on Confirm.' so failure is loud without blocking the user.
2. LLM telemetry — fallback was masking infra health. llmClient now
logs every outcome with elapsed time: 503/404/502/504/non-2xx/parse
failures all console.warn with reason + ms. Success path
console.info with provider + content length. Friendly fallback to
the user stays the same; ops/devs see the real story.
3. Catalog hygiene — duplicate 'Laptop Procurement' rows. list_processes
now dedupes by normalized display_name after the existing _key
dedupe. EA2 seeds + tenant imports both publishing the same name
collapse to one row in the assistant output.
4. Mission intuitiveness — user said 'I don't understand anything that's
going on there.' Added a one-line first-visit primer strip
('What you're looking at. Each tab below is a process running in
your company...') dismissible with an X, sticky to localStorage.
Also fixed the stale 'showing snapshot' fallback copy in the
liveError banner (snapshot mode was removed in wave 1; banner now
says 'last known state shown').
30/30 vitest pass. tsc + vite build green.
Three Oracle-flagged gaps:
1. Studio createDraft 500 (was masked by retries). Empirical curl
reproducer proved POST /api/ea2/flow returns 201 with the minimal
shape but intermittently 500s when config.wizard.* is in the
initial body. Split: POST creates the draft minimally, then a
follow-up PUT attaches the wizard config best-effort. Draft
creation now lands reliably.
2. LLM 'temporarily unavailable' was a dead end for the user. Now
the assistant returns ok:true with a concrete menu of tools the
user CAN use right now ('list processes', 'start <name>',
'open <hub>', 'tell <name> that ...', 'remember ...', 'recall ...').
The reply renders as a normal assistant message, not a red error.
3. Topbar UI/UX Pro Max pass:
- Added aria-label to the ⌘K and avatar buttons (icon-only a11y)
- Removed the duplicate topbar-mid chips on Mission (Mission scene
already shows family/defName/version inline). Cleaner topbar.
- Added focus-visible outlines (2px amber, 2px offset) on all
topbar action buttons + user-menu items
- Hover/active scale on the avatar (1.04 / 0.97) with 160ms ease
- Tabular numerals for the ⌘K kbd and notification badge
- prefers-reduced-motion respected
30/30 vitest pass. tsc + vite build green.
POST /api/ea2/flow currently returns 500 after 19s on demo upstream.
Retry once with backoff, then surface a clear 'EA2 backend temporarily
unable to create new processes' message naming the status. The user
gets actionable info instead of a raw 500.
- chatApi.listThreads: catch on the edges query so a fresh empty
inbox doesn't bubble a 500 to the UI. New users see 'No
conversations yet' instead of the error banner.
- agentTools.llmFallback: when the LLM gateway returns 503/500,
show a clear 'plain-language replies temporarily unavailable'
message naming the reason, listing the deterministic capabilities
that still work. No more silent null returns.
If buildLiveScenariosFromApi hangs or errors on any EA2 read,
loginAs would await it forever, leaving the user stuck on
'VERIFYING...' even after auth+identity succeed. Make refreshLive
fire-and-forget after the toast; the Mission scene re-fetches
on mount and surfaces failures via liveError.
CoreDNS at 10.43.0.10 returns SERVFAIL for demo.flow-master.ai in
this environment. Including it in the resolver list caused nginx
to query it first, get SERVFAIL, and return 502 before falling
back to 1.1.1.1. Pinning to public resolvers eliminates the cold
miss. 300s TTL caches the result.
Restored $variable form for proxy_pass because the static-hostname
form failed pod readiness (DNS query at config-load time hits
the same SERVFAIL).
The previous $variable form forced lazy per-request DNS resolution.
Per-request lookups against 10.43.0.10 → SERVFAIL → cold-cache 502
even when downstream was healthy. Switching to a literal hostname
makes nginx resolve once at startup and reuse the cached upstream IP
for the lifetime of the worker. Cloudflare anycast is stable so this
is safe.
Oracle round-9 watch-out: same liveMeta.fetchedFrom leak class
existed in two more user-visible surfaces:
- src/components/Telemetry.tsx: 'source' block rendered the host
string. Now hard-coded 'EA2'.
- src/scenes/Landing.tsx: hero paragraph said 'runs end-to-end
through EA2 on <host>'. Dropped the host suffix entirely.
liveMeta import removed from Telemetry (no other refs). Landing
still uses liveMeta for the workItems/distinctDefs stat counters.
Oracle round-9 third blocker: App.tsx topbar tag rendered
liveMeta.fetchedFrom?.replace('https://', '') which surfaced
'demo.flow-master.ai' in the topbar chip in live mode.
Replaced with the buyer-safe brand label 'live · EA2'. liveMeta
import removed (no other references). Wizard preview pane label
change and CSP recovery from prior commit stay.
When the K3s node recovers from an outage, cluster DNS (CoreDNS at
10.43.0.10) can be briefly unavailable, and nginx's startup-time
resolution of demo.flow-master.ai + canvas-llm-proxy.demo.svc...
fails the conf check, leaving the pod in CrashLoopBackOff. The
hostnames sat in proxy_pass directives that nginx resolves at boot.
Switched to:
- resolver 10.43.0.10 valid=30s ipv6=off;
- set $upstream_demo / $upstream_llm; proxy_pass $variable;
This forces lazy per-request DNS so transient cluster-DNS hiccups
don't kill the canvas-frontend pod. Verified during a real K3s node
outage today.
The seedHeadsIfMissing path I added fired N listMessages calls on
first refresh after auth, blowing idle-poll budget (landing went to
11 reqs at 30s). Reverting to the prior behavior:
- Topbar Chat badge shows real unread when localStorage heads are
populated (i.e. after the user has visited Chat at least once)
- Cold-start (fresh browser / cleared storage) the badge shows 0
until Chat is opened — same accepted caveat Oracle named on
sha-c3515707 and PASS'd
Idle-poll budget protected; Wizard 'Manual review' label change
from prior commit stays.
Close Oracle's two standing caveats so they can't gate a future loop:
1. Wizard dispatch label 'Person' was vague (Oracle: could imply chat
recipient, not work-routing). Now 'Manual review' in both dropdown
and preview-pane badge. dispatch_kind storage value unchanged
('human'/'agent'/'system' for EA2).
2. Chat unread cold-start. Topbar badge depended entirely on
localStorage being populated by Chat scene's first visit. Fresh
browser / private window / cleared storage saw 0 unread until then.
Added seedHeadsIfMissing in App.tsx topbar poll: on first refresh
(and only when localStorage heads count < thread count), fetches
last message per thread ONCE and persists to localStorage. Subsequent
ticks reuse the cache — same idle-poll budget (test still PASS).
Oracle narrow FAIL on sha-a02cc70e: Inspector.tsx:91 still emitted
'Switch to LIVE mode + sign in to execute "..." against
/api/runtime/transactions/{id}/actions/{actionId}.'
Rewritten as 'Switch to live mode and sign in to run "..." on this
case.' Matches the LeftRail and Approvals toast voice.
Source grep for user-visible /api/ in pushToast/title/JSX text returns
zero matches.
Oracle PASS-with-caveats; closing all 3:
1. Wizard dispatch labels were 'human'/'agent'/'system' (impl-ish).
Now 'Person'/'Assistant'/'System' in both the dropdown and the
preview-pane badge.
2. Chat badge was thread COUNT (a buyer with 5 chats sees '5' forever).
Now real unread count, computed locally from localStorage heads +
read maps that the Chat scene already persists. App.tsx topbar
poll still only fires 2 network calls (listThreads + workItems);
per-thread unread math is zero-network.
3. LeftRail toast leaked 'POST /api/runtime/transactions/{id}/actions/
submit' / save_draft. Rewritten as 'Switch to live mode and sign
in to submit this action against EA2.'
Plus engineering-string sweep elsewhere: 'demo.flow-master.ai' /
'bundled JSON' / 'in-browser fetch' references removed from:
- App.tsx mode toggle title
- MissionControl loading spinner
- Landing mode-button title
- state/store snapshot toast
- data/live.ts + data/synthetic.ts tour-step bodies
- buildScenarios.ts tagline
Source-grep confirms no remaining user-visible engineering jargon
referencing the backend hostname or raw endpoints.
The new topbar badge on the Chat tab (chatThreadCount > 0) broke the
old /^\\s*Chat\\s*$/ regex. Switched to a Playwright filter chain:
.tab filtered by hasText 'Chat' AND hasNotText 'Assistant' (which also
contains 'Chat' substring elsewhere in DOM). All 4 QA suites green:
- 36/36 dogfood (chat_sidebar threads=8 previews=8 times=8)
- 12/12 buyer-script (Approvals action disabled? true)
- 8/8 idle-poll (0 reqs at 30s on every scene)
- 14/14 mobile (390px clean)
User brief: 'normal businessman could use this'. The command palette
was already wired but leaked engineering jargon and missed half the
scenes.
- 'Scenes' renamed 'Go to'; now lists all 14 user-reachable scenes
(was 5). Hubs, Documents, Approvals, Chat, Assistant, Explainer,
Geo - all keyboard-reachable.
- Endpoint strings ('POST /api/runtime/transactions', 'demo.flow-
master.ai', 'bundled JSON') removed from hints. Replaced with
buyer-safe equivalents like 'writes to EA2' / 'for developers'.
- Real-actions group: hints simplified, less raw IDs.
- Dropped 'Dispatch sidekick agent · coming soon' (dead stub).
- 'Data mode' group renamed 'Preferences' and simplified to 3
options (refresh live, theme toggle, dev console).
Idle-poll audit caught a regression I introduced in the prior commit:
the topbar badge poll iterated chatApi.listMessages per thread on each
60s tick, firing N requests per refresh. Landing scene showed 9 reqs
in a 30s idle window (failed the <=5 budget).
Now the topbar badge just shows the THREAD COUNT (not per-thread
unread), which is a single chatApi.listThreads call. Per-thread unread
math stays in the Chat scene where it belongs (already wired:
chat-thread-row.unread + .chat-unread-summary).