Oracle wave-4: 'Re-test LLM with a successful upstream response,
not only the 500/503 fallback path. The current evidence proves
graceful degradation but not that LLM is actually on producing
model-backed answers.'
The user said 'Take the LLM integration that we already have on
the other websites and put it on.' demo's llm-integration-service
is 503 (no provider key), but dev.flow-master.ai's identical
gateway returns real OpenAI-backed completions and accepts the
canvas-issued JWT plus has CORS open to canvas.flow-master.ai.
Refactor llmClient:
- Extract callGateway() that returns LlmResponse or 'retry' sentinel
- LLM_GATEWAYS = [primary (canvas's own), fallback-dev (dev.fm.ai)]
- Retry on 500/502/503/504/404 and network errors against the next gateway
- Telemetry prefixes each log with gateway name ('primary' vs 'fallback-dev')
- 12s deadline still applies across the whole cascade
- Only final 'all gateways exhausted' warn if every link in the chain fails
Net result: canvas users get a real model-backed answer on every
non-tool prompt instead of the 'temporarily unavailable' fallback.
Oracle wave-3 caveat: 'A toast alone is insufficient if the app marks
the wizard state as attached locally and never retries.'
Track draft.wizardConfigAttached. Set true only after the initial PUT
resolves. In handleStructureSave (Confirm Structure), if not yet
attached, retry updateDraftConfig before applying the structure batch.
Failure of the retry is logged but doesn't block save — structure
still lands and downstream config sync can pick up later.
Bug: 'tell me a haiku about flowcharts' matched send_chat with
target='me a haiku' / body='flowcharts', producing the bogus
"I don't know who 'me a haiku' is" reply and silently swallowing
all non-tool prompts that started with tell/ask. The LLM fallback
never ran, so [llm] telemetry from wave 3 never fired in tests.
Fix: tighten the matcher to require a single-word [A-Za-z]+ name
between the verb and that/about/:. Drop the substring search in the
directory; require an exact match. Now 'tell me a haiku about X'
falls through to llmFallback as intended.
Investigation: the wave-3 [llm] telemetry warns weren't firing on live
because the fetch had no client-side timeout. When demo's
llm-integration-service was slow/stuck, fetch hung indefinitely past
the test's 12s wait, so the fallback path never ran and the user saw
no agent reply at all.
Add a 12s AbortController-backed deadline. On timeout: console.warn
'[llm] gateway timed out after Xms (deadline 12000ms)' + return the
graceful 'temporarily unavailable' fallback that renders the tool
menu in the chat surface. Uses AbortSignal.any when available to
preserve any caller-supplied signal.
Oracle's wave-2 watch-outs:
1. Wizard config persistence — the follow-up PUT was silently swallowed.
Empirical curl reproducer confirmed config.wizard DOES land after the
split create/PUT (config.wizard.marker=EA2_DRAFT_PROCESS visible on
re-GET). But to satisfy the 'no silent loss' concern: the catch now
warns to console AND surfaces a soft toast 'Draft saved, but its
wizard state didn't attach. You can keep working — save will retry
on Confirm.' so failure is loud without blocking the user.
2. LLM telemetry — fallback was masking infra health. llmClient now
logs every outcome with elapsed time: 503/404/502/504/non-2xx/parse
failures all console.warn with reason + ms. Success path
console.info with provider + content length. Friendly fallback to
the user stays the same; ops/devs see the real story.
3. Catalog hygiene — duplicate 'Laptop Procurement' rows. list_processes
now dedupes by normalized display_name after the existing _key
dedupe. EA2 seeds + tenant imports both publishing the same name
collapse to one row in the assistant output.
4. Mission intuitiveness — user said 'I don't understand anything that's
going on there.' Added a one-line first-visit primer strip
('What you're looking at. Each tab below is a process running in
your company...') dismissible with an X, sticky to localStorage.
Also fixed the stale 'showing snapshot' fallback copy in the
liveError banner (snapshot mode was removed in wave 1; banner now
says 'last known state shown').
30/30 vitest pass. tsc + vite build green.
Three Oracle-flagged gaps:
1. Studio createDraft 500 (was masked by retries). Empirical curl
reproducer proved POST /api/ea2/flow returns 201 with the minimal
shape but intermittently 500s when config.wizard.* is in the
initial body. Split: POST creates the draft minimally, then a
follow-up PUT attaches the wizard config best-effort. Draft
creation now lands reliably.
2. LLM 'temporarily unavailable' was a dead end for the user. Now
the assistant returns ok:true with a concrete menu of tools the
user CAN use right now ('list processes', 'start <name>',
'open <hub>', 'tell <name> that ...', 'remember ...', 'recall ...').
The reply renders as a normal assistant message, not a red error.
3. Topbar UI/UX Pro Max pass:
- Added aria-label to the ⌘K and avatar buttons (icon-only a11y)
- Removed the duplicate topbar-mid chips on Mission (Mission scene
already shows family/defName/version inline). Cleaner topbar.
- Added focus-visible outlines (2px amber, 2px offset) on all
topbar action buttons + user-menu items
- Hover/active scale on the avatar (1.04 / 0.97) with 160ms ease
- Tabular numerals for the ⌘K kbd and notification badge
- prefers-reduced-motion respected
30/30 vitest pass. tsc + vite build green.
POST /api/ea2/flow currently returns 500 after 19s on demo upstream.
Retry once with backoff, then surface a clear 'EA2 backend temporarily
unable to create new processes' message naming the status. The user
gets actionable info instead of a raw 500.
- chatApi.listThreads: catch on the edges query so a fresh empty
inbox doesn't bubble a 500 to the UI. New users see 'No
conversations yet' instead of the error banner.
- agentTools.llmFallback: when the LLM gateway returns 503/500,
show a clear 'plain-language replies temporarily unavailable'
message naming the reason, listing the deterministic capabilities
that still work. No more silent null returns.
If buildLiveScenariosFromApi hangs or errors on any EA2 read,
loginAs would await it forever, leaving the user stuck on
'VERIFYING...' even after auth+identity succeed. Make refreshLive
fire-and-forget after the toast; the Mission scene re-fetches
on mount and surfaces failures via liveError.
The previous $variable form forced lazy per-request DNS resolution.
Per-request lookups against 10.43.0.10 → SERVFAIL → cold-cache 502
even when downstream was healthy. Switching to a literal hostname
makes nginx resolve once at startup and reuse the cached upstream IP
for the lifetime of the worker. Cloudflare anycast is stable so this
is safe.
Oracle round-9 watch-out: same liveMeta.fetchedFrom leak class
existed in two more user-visible surfaces:
- src/components/Telemetry.tsx: 'source' block rendered the host
string. Now hard-coded 'EA2'.
- src/scenes/Landing.tsx: hero paragraph said 'runs end-to-end
through EA2 on <host>'. Dropped the host suffix entirely.
liveMeta import removed from Telemetry (no other refs). Landing
still uses liveMeta for the workItems/distinctDefs stat counters.
The new topbar badge on the Chat tab (chatThreadCount > 0) broke the
old /^\\s*Chat\\s*$/ regex. Switched to a Playwright filter chain:
.tab filtered by hasText 'Chat' AND hasNotText 'Assistant' (which also
contains 'Chat' substring elsewhere in DOM). All 4 QA suites green:
- 36/36 dogfood (chat_sidebar threads=8 previews=8 times=8)
- 12/12 buyer-script (Approvals action disabled? true)
- 8/8 idle-poll (0 reqs at 30s on every scene)
- 14/14 mobile (390px clean)
Idle-poll audit caught a regression I introduced in the prior commit:
the topbar badge poll iterated chatApi.listMessages per thread on each
60s tick, firing N requests per refresh. Landing scene showed 9 reqs
in a 30s idle window (failed the <=5 budget).
Now the topbar badge just shows the THREAD COUNT (not per-thread
unread), which is a single chatApi.listThreads call. Per-thread unread
math stays in the Chat scene where it belongs (already wired:
chat-thread-row.unread + .chat-unread-summary).
When the QA runs against a fresh CEO session (no pre-existing chat
threads), the new-thread create + inbox-edge + message send chain is
4-8 EA2 round-trips. Prior waits (8s thread + 2.5s send + 12s
condition) were occasionally too tight. Bumped to 15s + 4s + 20s.
36/36 dogfood stable across the slow-path.
Buyer-script audit showed queue row titles falling through to raw
case-key hex when business_subject was empty / set to the case key
itself. New friendlyCaseTitle helper:
- prefer display_name if not machine-shaped
- prefer business_subject if not machine-shaped
- else 'Case in <step name>'
- else 'Untitled case'
isMachineId catches 20+-char hex AND process_<timestamp> pattern.
Buyer-script QA caught two regressions:
1. Toast on action failure leaked 'runtime-values isn't reachable' and
'Backend can't accept' — engineering jargon a buyer would read as
'product is broken'.
2. After a real 500, the button stayed enabled so a buyer could click
again into the same error.
Now:
- Failure toast says 'The runtime engine couldn't accept that decision.
Action buttons are now disabled until it recovers.' (no jargon, no
raw error strings).
- handleAction sets runtimeHealth='down' on the failure, which flips
the runtime-note banner on and disables every action button.
- Non-runtime failures get the even-shorter 'Something went wrong
handling that action. Try again in a moment.'
Closes the 'fork Pi coding agent' integration loop with a tiny
broker that mirrors the @earendil-works/pi-ai request shape:
POST /internal/canvas-llm/chat
body: { messages: [{role, content}], max_tokens }
The proxy speaks to OpenAI or Anthropic depending on which env key
is set. With no key, returns 503 so the frontend's deterministic
fallback in routeAgentInput fires gracefully.
- llm-proxy/main.py: FastAPI + httpx (~150 LOC). System messages get
split out for Anthropic (separate 'system' field) and inlined for
OpenAI.
- llm-proxy/Dockerfile: python:3.12-slim, uvicorn on :8080.
- Deployment + Service in demo namespace (verified Running).
- nginx.conf: /internal/canvas-llm/ proxies to the in-cluster service.
Provider keys are set to empty in the manifest by design: customer
flip is a single 'kubectl set env' or sealed-secret update.
Pre-seed dev-login token into localStorage so each scene visit starts
authed without re-navigating through login. Use fresh p.goto per scene
instead of p.goBack so the viewport stays pinned to 390x844.
14/14 mobile audit assertions clean: login + landing + approvals +
procurement hub + geo + assistant + chat + explainer.
Mobile audit at 390px viewport showed every scene with topbar
expanded to ~1333px because of fixed-width brand-lock + tabs +
topbar-actions in a grid-template-columns auto auto 1fr auto.
- Topbar collapses to single column with scrollable tab row
- Tabs become horizontally scrollable; smaller font/padding
- Hide topbar-mid (context chips) and user-email/topbar-age (the
user knows who they are)
- Approvals split collapses at 700px instead of 900px
- Compact scene padding (16/12) on mobile
Plumbed me.tenant_id through api.ping + loginAs into the store as
tenantId. Approvals.loadQueue filters /api/ea2/work-items by
tenant_id matching the signed-in user so the queue only shows
actionable rows. Probe of EA2 work-items returned 73/80 in canvas's
tenant (a0000000-...), 6 in hms_dev, 1 in hub-cnt-f56ce0 -- those
7 are now hidden.
The strict QA gate requires timeBadges === threadRowCount. When a
thread has no preview message (e.g. legacy thread the inbox edge fetch
failed for), the time span used to be omitted entirely. Now renders
'—' for missing preview so the per-row time element exists.
Oracle round-9 hard blocker: live CSP advertised script-src 'self'
'unsafe-inline' 'unsafe-eval'. Verified the Vite build emits zero inline
<script> tags (only external module script) and no bundled dep uses
new Function / eval (reactflow / leaflet / zustand / dagre / cmdk /
framer-motion all clean).
- nginx.conf script-src reduced to 'self'.
- img-src extended for OSM tiles (https://*.tile.openstreetmap.org) and
leaflet marker images served from unpkg.
Also (Oracle round-9 second blocker) chat_sidebar QA now creates a
thread + sends a message before asserting, then strict-asserts
threadRowCount >= 1 AND previewRows === threadRowCount AND
timeBadges === threadRowCount. No more vacuous PASS when sidebar is
empty.
Oracle round-8 closure: 3 user-visible 'Pi' refs remaining in
Explainer.tsx ('ask Pi' in DEFINE step, 'Why an agent (Pi)?' card
title, 'Pi is a co-pilot' card body, 'Ask Pi instead' chip).
All replaced with 'the assistant' / 'Command Assistant'. New
no_pi_branding_in_user_facing_scenes QA gate sweeps landing +
explainer + agent for \bPi\b word boundary on the rendered DOM
text and fails if any scene leaks.
Three Oracle round-7 caveats:
1. Wizard view replacement now deletes old presentation edges before
creating new field-based ones. Previously handleDataSave added new
edges additively, so runtime could still hit the generic Notes view.
2. Chat sidebar QA assertion was vacuously true (previewRows >= 0 &&
timeBadges >= 0). Now: when threads exist, preview count must equal
thread count AND at least one relative timestamp must render.
3. Agent composer placeholder rebranded from 'Ask Pi...' to 'Ask the
assistant...'. Closes residual Pi/LLM framing leak.
Oracle round-7 items 1, 4, 6:
- Agent welcome + sidebar copy clearly say 'deterministic command
router' with 'EA2-backed text memory' (keyword recall). LLM adapter
described as 'separate component, off by default'. No more 'natural
language LLM on the roadmap' overclaim.
- 3 new full_dogfood gates:
* memory_vault_remember_acks (proves vault write)
* memory_vault_recall_returns_match (proves vault read by content)
* llm_unconfigured_falls_back_gracefully (proves agent never crashes
when LLM proxy isn't wired - the current state)
* chat_sidebar_shows_previews_and_timestamps (proves chat polish
elements render)
Assistant grew remember + recall tools this loop (7 total). The previous
exact-count assertion is now an at-least check so future tool additions
don't break QA.
src/lib/llmClient.ts: browser POSTs to /internal/canvas-llm/chat with
{messages, max_tokens}. Mirrors the normalized request shape used by
earendil-works/pi's @earendil-works/pi-ai. Returns:
- ok:true with content+provider when configured and successful
- ok:false, configured:false when proxy 404/503/unreachable
- ok:false, configured:true when provider returned an error
routeAgentInput now falls back to the LLM with a system prompt that
includes the user identity, tool catalogue, and the top-3 vault recall
hits when no deterministic tool matches. When the proxy is not
configured the deterministic vault-hint reply is used (no regression
in the demo tier).
The /api/ea2/flow/processes endpoint caps at ~203 items and pr_to_po_def
is not in that window for this tenant despite being status=published.
fetchStartableFlows() does a direct GET per allowlist key and merges
results with the catalogue list (allowlist hits take precedence).
Wired into Hub.loadPublishedFlows + agentTools.list_processes +
agentTools.start_process.
Previous attempt blacklisted unstartable backfill source_contexts but
swept all real published flows out too. Inverted to a positive
allowlist: only flows whose _key is in STARTABLE_FLOW_KEYS reach the
Hubs / Assistant.
Today the canonical pr_to_po_def is the only runtime-startable flow in
the tenant. New entries are added by hand once we've verified a flow
actually starts via /api/runtime/transactions. Cleaner than chasing
source_context strings.
Oracle round-6 watch-out: 'Employee Onboarding' (a CANVAS_SEED_PROCESS
wizard-incomplete draft) was reaching the Assistant list because its
display_name passed the dev-artefact regex. EA2 runtime then returned
500 transaction_creation_failed because it lacked view/data attachments.
- flowCuration.NON_BUSINESS_SOURCE_CONTEXTS now drops CANVAS_SEED_PROCESS,
CANVAS_SEED_STEP, CANVAS_CHAT_INBOX, CANVAS_ATTENDANCE_ROOT,
CANVAS_ATTENDANCE_SITE - all internal canvas source_contexts.
- qa/full_dogfood.mjs new assertion 'assistant_list_processes_all_startable'
posts each Assistant-listed flow to /api/runtime/transactions and
asserts non-4xx/5xx response.
Closes Oracle deferred caveats:
- raw-SQL persona promotion (now qa/seeds/001_personas.mjs, idempotent)
- no scripted multi-persona interaction (now qa/seeds/002_persona_
interactions.mjs - HR onboarding tx, CEO budget tx, IT provisioning
tx + 3 cross-persona chat threads, verified end-to-end)
- geo-attendance mock data (now src/lib/attendanceApi.ts pulls 8 sites
via flow.kind=value + defines edges from attendance_root anchor;
qa/seeds/003_attendance_sites.mjs seeds them idempotently)
Each site is a flow.kind=value doc with config.attendance = { lat,
lng, kind, status, who, city, label }. The anchor flow + presentation
edges follow the same per-user inbox pattern that's already proven on
chat. No hardcoded JS array left in src/scenes/GeoAttendance.tsx.
Single QA script (qa/full_dogfood.mjs) drives the live URL like a real
operator and asserts all of Round 2:
auth_guard_unauthed_redirects_to_login
login_shows_three_personas
login_has_microsoft_sso_button
ceo_persona_dev_login_lands_on_landing
landing_exposes_all_hubs_and_extras (7/7)
explainer_renders_eight_cards
geo_attendance_renders_tiles + 8 markers
pi_agent_shows_five_tools
pi_agent_list_processes_returns_real_ea2_data
procurement_hub_lists_real_flows
chat_scene_loads_with_thread_sidebar
theme_toggle_flips_theme
security_header_* x6 (HSTS, X-Frame-Options, X-CTO, Referrer-Policy,
CSP, Permissions-Policy)
Exits non-zero on failure so the script is CI-grade.
Eight-card explainer covering Business-as-Code philosophy, the five
element types (Flow/Data/View/Rule/Version), the DEFINE/VERSION/DEPLOY/
EXECUTE lifecycle, hubs, work items, EA2 graph storage, and the Pi
agent. Sourced from FlowMaster Element Architecture v2.1 + the EA2
collection/edge contracts recalled from Hindsight today.
New 'What is FlowMaster?' chip on Landing routes here.
- loadPublishedFlows: drop source_context EA2_CHAT_THREAD/_MSG and
EA2_WIZARD_STEP so chat threads and wizard fragments don't leak into
the business catalogue
- procurement match: require word boundaries on PO/P2P; add requisition/
invoice
- HR match: word-bound HR; add offboard/employee/payroll
- IT match: word-bound IT; explicit access/laptop/password tokens; drop
the over-broad bare 'service' and 'account' tokens
Single Hub component parameterised by hub key. Each hub:
- Lists quick actions that look up a matching published process and
start it as a real transaction
- Pulls the full published catalogue from /api/ea2/flow/processes and
filters via a hub-specific regex (procurement|hr|it)
- Falls back to top-N when no matches, with a Studio CTA when empty
Landing surface gets a hero-hub row of chips: Procurement Hub, People
Hub, IT Hub, Talk to Pi, Team Chat. Same EA2-write doctrine as wizard.
process-definitions/{key}/graph filters to published process definitions
and returns empty for chat-thread flows, so the second persona saw no
message bodies. The raw defines edge endpoint returns the link correctly.
New Chat scene at the Chat top-bar tab. Threads and messages are stored
as flow docs (kind=definition, source_context=EA2_CHAT_THREAD / _MSG)
linked by defines edges with role=next — same proven contract as wizard
steps. Persona directory lets the signed-in operator open a thread with
the CEO / HR / IT personas.
- src/lib/chatApi.ts: listThreads / createThread / listMessages / sendMessage
- src/scenes/Chat.tsx: sidebar (threads + new-thread picker) + main pane
(header, bubbles, composer with Enter-to-send / Shift-Enter newline)
- src/index.css: chat-scene grid + bubble + composer styles
- src/App.tsx: Chat tab + scene route
Backend EA2 apply-batch requires a request_id (idempotency key). Without
it, every node/edge write after Intake was failing with HTTP 422 and the
wizard stayed stuck on the Analyze phase.
Verified via qa/audit_wizard_writes.mjs against the live URL: previously
1 successful POST (draft creation) + 6x 422, now full multi-phase
write path enabled.
Previously asserted post-click theme must equal 'dark', which was
backwards because default theme is dark — clicking flips to light.
Now record before, click, record after, assert they differ.
Also tighten the selector to .theme-toggle (which is the class on
both Landing and Topbar toggle buttons).
Two bugs:
1. [data-theme="dark"] block was defined BEFORE :root in source CSS,
so :root cascade overrode it. Moved to AFTER :root and bumped
specificity to html[data-theme="dark"] to beat plain :root.
2. Theme toggle button only existed in the topbar, which is hidden
on landing/login/sso-callback scenes. Added a ThemeToggle to
the Landing header so users can flip theme from the home page.
Confidence: high
Scope-risk: narrow
The Vite minifier was tree-shaking the dev-login button because
devLoginConfig() always returned {enabled: false} when the
/internal/dev-login-config endpoint 404'd (which is the case on
canvas, where the endpoint isn't mounted).
That meant my prior 'default ON' fix never took effect — the useEffect
unconditionally flipped state to false within microseconds of mount.
Real fix: return {enabled: null} on absence. Only flip OFF when
endpoint explicitly says enabled=false. Otherwise honor the source
default.
Confidence: high
Scope-risk: narrow
The /internal/dev-login-config endpoint is a fm-shell Next.js internal
route that doesn't exist on the canvas nginx. Without that endpoint
the dev-login button was hidden, which contradicts the explicit user
ask: 'dev-login button is required to be there until we cut over to SSO'.
Fix: default state to true, treat the config endpoint as 'flip OFF only'.
If the endpoint returns enabled:false we hide. Otherwise (404, network
error, etc.) we keep the button visible.
Confidence: high
Scope-risk: narrow
Not-tested: real /internal/dev-login-config returning enabled:false (would
need that endpoint mounted on canvas — out of scope for now)