FieldWorrk collapses requirements-gathering, architecture, and build into one AI-guided conversation. My job as lead designer wasn't making the AI smarter — it was making a risk-averse enterprise user comfortable enough to click "Approve" on what it produced.
Leadership's ask was blunt: enterprise clients want to go from business idea to working app without waiting six weeks for a requirements document. Can AI collapse that timeline — and can we make people trust it enough to actually use it?
The second half is what this case study is about. Building an AI that drafts a data model in nine seconds is a research problem. Getting a risk-averse ops manager to click "Approve" on that model without personally re-deriving it is a design problem.
High-fidelity, working recreations of the four screens — not screenshots. Try them the way the research below describes them being used.
This is the round-2 fix: five falsifiable fields instead of one AI-written paragraph. Editing one doesn't touch the others.
I've updated the process with the sub-steps and automated notifications you described. Compare both versions below.
Same process, before and after. A process owner sees exactly what the AI proposes to automate before approving anything.
Every expansion is where a judgment call could hide. That's what the assumptions counter is tracking as each step gets designed.
Errors explain themselves in place and offer a next action — no separate error console to go find.
Before touching a screen, I ran 14 contextual interviews with business analysts, process owners, and citizen developers across our banking and retail pilot accounts, focused on how they currently move a process from idea to deployed system, and where it breaks.
By the time I've written the requirements doc, had it reviewed, and handed it to the dev team, the business has usually changed their mind about half of it.— P3, Business Analyst, Banking
I don't trust a system that just tells me "done." I need to know what it assumed, because I'm the one who gets asked questions six months later when something breaks.— P7, Process Owner, Retail Ops
Every low-code tool I've used either treats me like I can't understand the technical side at all, or dumps raw JSON on me. There's no middle ground.— P11, Citizen Developer
Three principles came out of this and shaped every decision after:
Never say "done" without saying based on what. The system has to show its reasoning, not just its output.
Keep the conversation and the record of truth separate. No one should have to re-read chat scrollback to know what was decided.
Meet people at their actual technical level — readable summaries by default, technical detail on demand, never the reverse.
The first prototype covered only the Discover phase: chat on the left, one free-text AI summary on the right. We ran a moderated test with 8 participants, each describing a real process from their own job.
Three participants independently said some version of not being able to tell what the AI actually understood versus what it was just repeating back.
This just looks like it rephrased what I said. I don't know if it actually extracted anything or if it's just... summarizing.— P4, usability session
This was the most important finding of the project. The problem was never the AI's accuracy — it was legibility of its reasoning. Users couldn't distinguish "the AI understood this" from "the AI echoed this."
We broke the free-text summary into a Discovery Snapshot — discrete, independently editable fields. Not a visual redesign so much as an epistemic one: forcing the AI's understanding into named slots a user could individually confirm or correct.
P7's comment about being "the one who gets asked questions six months later" is a liability problem wearing a UX problem's clothes. So every time the AI made a judgment call rather than working from explicit input, it logged it — and the running count surfaced right in the chat.
Tested with 6 participants in governance-heavy roles: 5 of 6 said the visible count alone increased their willingness to approve a design without opening every assumption. But one gap surfaced too:
I want to trust the number, but numbers I can't check aren't audits, they're just... vibes.— P2, usability sessionShipped without drill-down — prioritized next
Same task, 8 returning plus 4 new participants (n=12), three weeks later.
Median time to a positive trust comment moved from 3:40 (negative sentiment) to 6:55 — and the sentiment flipped from suspicion to "oh, I can just fix this field."
Users didn't need the AI to be more accurate. They needed a surface for disagreement. Editable, discrete fields gave them that surface.
Six weeks into general availability, we pulled a phase-completion funnel across every solution started by early-access clients.
Build also had 3.4× longer average dwell time than any other phase. Qualitatively, this was where the interaction paradigm silently switched from conversation to dashboard, with no transition to prepare users for it.
Build is fundamentally an async, multi-system status board, and that doesn't finish in the same conversational rhythm as Discover or Reimagine. I don't think we designed the seam between Design and Build carefully enough.— Engineering Lead, project retro
Every phase pairs an ephemeral chat with a persistent, structured record. When we asked participants to find a decision made moments earlier, 12 of 12 went straight to the structured panel, not the scrollback — even though the decision was made in the chat. That's strong evidence the split matches how people actually think about "discussed" vs. "decided."
Given the funnel data and the retro comment above, I'd insist on a small transitional moment between Design and Build — even one message reframing the shift. We treated it as a visual polish problem when it was really a mental-model transition, and no amount of icon or color work fixes that.
Assumptions registry shows a count but no drill-down — P2's "vibes, not audits" feedback is still open, now top of next quarter's backlog.
No onboarding moment at the Design → Build transition, despite clear funnel and qualitative evidence it's needed.
Status vocabulary is inconsistent across phases — flagged by our own heuristic review, not yet studied with real users.