Bug: 'tell me a haiku about flowcharts' matched send_chat with
target='me a haiku' / body='flowcharts', producing the bogus
"I don't know who 'me a haiku' is" reply and silently swallowing
all non-tool prompts that started with tell/ask. The LLM fallback
never ran, so [llm] telemetry from wave 3 never fired in tests.
Fix: tighten the matcher to require a single-word [A-Za-z]+ name
between the verb and that/about/:. Drop the substring search in the
directory; require an exact match. Now 'tell me a haiku about X'
falls through to llmFallback as intended.
Investigation: the wave-3 [llm] telemetry warns weren't firing on live
because the fetch had no client-side timeout. When demo's
llm-integration-service was slow/stuck, fetch hung indefinitely past
the test's 12s wait, so the fallback path never ran and the user saw
no agent reply at all.
Add a 12s AbortController-backed deadline. On timeout: console.warn
'[llm] gateway timed out after Xms (deadline 12000ms)' + return the
graceful 'temporarily unavailable' fallback that renders the tool
menu in the chat surface. Uses AbortSignal.any when available to
preserve any caller-supplied signal.
Oracle's wave-2 watch-outs:
1. Wizard config persistence — the follow-up PUT was silently swallowed.
Empirical curl reproducer confirmed config.wizard DOES land after the
split create/PUT (config.wizard.marker=EA2_DRAFT_PROCESS visible on
re-GET). But to satisfy the 'no silent loss' concern: the catch now
warns to console AND surfaces a soft toast 'Draft saved, but its
wizard state didn't attach. You can keep working — save will retry
on Confirm.' so failure is loud without blocking the user.
2. LLM telemetry — fallback was masking infra health. llmClient now
logs every outcome with elapsed time: 503/404/502/504/non-2xx/parse
failures all console.warn with reason + ms. Success path
console.info with provider + content length. Friendly fallback to
the user stays the same; ops/devs see the real story.
3. Catalog hygiene — duplicate 'Laptop Procurement' rows. list_processes
now dedupes by normalized display_name after the existing _key
dedupe. EA2 seeds + tenant imports both publishing the same name
collapse to one row in the assistant output.
4. Mission intuitiveness — user said 'I don't understand anything that's
going on there.' Added a one-line first-visit primer strip
('What you're looking at. Each tab below is a process running in
your company...') dismissible with an X, sticky to localStorage.
Also fixed the stale 'showing snapshot' fallback copy in the
liveError banner (snapshot mode was removed in wave 1; banner now
says 'last known state shown').
30/30 vitest pass. tsc + vite build green.
Three Oracle-flagged gaps:
1. Studio createDraft 500 (was masked by retries). Empirical curl
reproducer proved POST /api/ea2/flow returns 201 with the minimal
shape but intermittently 500s when config.wizard.* is in the
initial body. Split: POST creates the draft minimally, then a
follow-up PUT attaches the wizard config best-effort. Draft
creation now lands reliably.
2. LLM 'temporarily unavailable' was a dead end for the user. Now
the assistant returns ok:true with a concrete menu of tools the
user CAN use right now ('list processes', 'start <name>',
'open <hub>', 'tell <name> that ...', 'remember ...', 'recall ...').
The reply renders as a normal assistant message, not a red error.
3. Topbar UI/UX Pro Max pass:
- Added aria-label to the ⌘K and avatar buttons (icon-only a11y)
- Removed the duplicate topbar-mid chips on Mission (Mission scene
already shows family/defName/version inline). Cleaner topbar.
- Added focus-visible outlines (2px amber, 2px offset) on all
topbar action buttons + user-menu items
- Hover/active scale on the avatar (1.04 / 0.97) with 160ms ease
- Tabular numerals for the ⌘K kbd and notification badge
- prefers-reduced-motion respected
30/30 vitest pass. tsc + vite build green.
POST /api/ea2/flow currently returns 500 after 19s on demo upstream.
Retry once with backoff, then surface a clear 'EA2 backend temporarily
unable to create new processes' message naming the status. The user
gets actionable info instead of a raw 500.
- chatApi.listThreads: catch on the edges query so a fresh empty
inbox doesn't bubble a 500 to the UI. New users see 'No
conversations yet' instead of the error banner.
- agentTools.llmFallback: when the LLM gateway returns 503/500,
show a clear 'plain-language replies temporarily unavailable'
message naming the reason, listing the deterministic capabilities
that still work. No more silent null returns.
If buildLiveScenariosFromApi hangs or errors on any EA2 read,
loginAs would await it forever, leaving the user stuck on
'VERIFYING...' even after auth+identity succeed. Make refreshLive
fire-and-forget after the toast; the Mission scene re-fetches
on mount and surfaces failures via liveError.
CoreDNS at 10.43.0.10 returns SERVFAIL for demo.flow-master.ai in
this environment. Including it in the resolver list caused nginx
to query it first, get SERVFAIL, and return 502 before falling
back to 1.1.1.1. Pinning to public resolvers eliminates the cold
miss. 300s TTL caches the result.
Restored $variable form for proxy_pass because the static-hostname
form failed pod readiness (DNS query at config-load time hits
the same SERVFAIL).
The previous $variable form forced lazy per-request DNS resolution.
Per-request lookups against 10.43.0.10 → SERVFAIL → cold-cache 502
even when downstream was healthy. Switching to a literal hostname
makes nginx resolve once at startup and reuse the cached upstream IP
for the lifetime of the worker. Cloudflare anycast is stable so this
is safe.