VX Development Blog

Retrieval-Augmented Workflow

Overview

Automation is at the heart of the software we design. If a task can be automated, we look for a way to automate it. Our flagship product Mulberry is no exception in this matter. Along with a long list of sophisticated features available to users, the system includes several automation tools like "automatic actions", scheduled jobs, maintenance tasks and more. Those tools can be configured, reconfigured and repurposed both by our clients and via the vast number of plugins developed for Mulberry.

The advancement of large language models and their availability through both proprietary APIs and open weight models opened a new frontier for our automations. It gave us the opportunity for "smarter" suggestions, targeted actions and safeguards against potentially wrong decisions. Though modern models can be very useful in our applications, one rule stays fixed: the final decision always belongs to the human operator.

The Design

New assistive tools in Mulberry are intended to help users during decision making by providing insights or suggestions. For example, there are tools for summarizing a user's inbox or a task's metadata along with its attachments and tools for suggesting an action or a recipient.

Analysis showed that all those tools share a common property: they all have deterministic workflows. Each tool performs the same checks and evaluates the same conditions a user would when working manually. Before any LLM inference, a tool runs a long series of these checks, any of which may terminate the process early or lead to a fixed outcome. Thus those tools are not AI agents, and the assistant is not agentic.

Still, these tools are not simple prompts either. They gather data from several sources, retrieve historical context, analyze attachments and return structured results. When we looked for an established term for this kind of capability, we found none that fit: too complex for a prompt, too structured for plain RAG, too constrained to be an agent. The closest fit was "workflow", which is accurate but too broad to be useful.

So we started calling them retrieval-augmented workflows.

A Retrieval-Augmented Workflow (RAW) is an application-controlled workflow that embeds AI reasoning into a predefined business process. The workflow owns execution control, context preparation, validation, and integration, while AI models perform bounded reasoning tasks such as classification, extraction, ranking, summarization, or recommendation.

RAG describes how knowledge is supplied to reasoning. RAW describes how AI reasoning is embedded into a business capability.

Retrieval-Augmented Workflow

Every RAW in Mulberry shares the same six properties.

Deterministic control flow

The execution sequence is written by developers, not chosen by the model. A RAW may branch and it may terminate early, but every possible path through it exists in the source code. The model reasons inside a step; it never decides which step comes next.

Context aggregation

Before any inference, a RAW collects the information a user would look at themselves: task metadata, attachments, the sender's history, the set of actions actually available in the current state, etc. Much of this is ordinary application code including permission checks, database queries, file parsing. Most of a RAW's implementation is exactly that.

Retrieval-augmented reasoning

Relevant past decisions are retrieved and passed to the model alongside the current case, for example, similar tasks, how they were handled and which workflows exist for this kind of work. This is the part a RAW shares with RAG.

Structured output

A RAW returns data, not conversation:

{
  "workflow_id": 42,
  "confidence": 0.91,
  "reasons": [
    "Matches sender category",
    "Similar historical submissions"
  ]
}

The shape matters. A suggestion the operator can't evaluate is a suggestion they have to trust blindly, so every result carries the reasons behind it and how confident the model is. The application validates this against a schema before it reaches the interface.

Invocation

A RAW is a capability, not an interface element. The same workflow runs when a user opens the assistant and invokes a tool, and when Mulberry's existing automation calls it via an automatic action, a scheduled job or a plugin. The workflow does not know or care which of the two triggered it: it receives a business request and returns a structured result either way.

This is a direct consequence of the deterministic control flow. A workflow with a fixed graph and a schema-validated result behaves identically whether a person is watching or not, which is what makes unattended execution reasonable in the first place.

Model-agnostic inference

The LLM call itself is a plugin point. A RAW declares what it needs (a prompt and the schema it expects back) and the actual inference is performed by a configurable backend: a proprietary API, a self-hosted open-weight model, or any other provider a client prefers.

This falls out of the same property as everything else here. Because the workflow owns the control flow, the model is a replaceable component inside one step rather than the thing running the show. Swapping providers changes where inference happens, not what the tool does.

Not an agent

A common assumption is that any multi-step AI feature is an agent. The distinction is not complexity but who owns the execution graph. In a RAW, the developer defines the workflow and the model performs reasoning within it. In an agent, the model defines the workflow as it goes. A system becomes agentic when it decides what data to retrieve, which tools to call, whether it needs more information, or how to revise its own plan.

This is a different axis from complexity, not a point further along the same one. A RAW can aggregate more data and run more steps than a simple agent and still not be agentic, because none of that changes who is driving.

The result

We have built and deployed dozens of RAWs in Mulberry, in contexts ranging from the inbox and task view to the action screen and submission registration (fig. 1). The architecture has since been ported to our other systems as well.

The Mulberry assistant

Fig. 1. The RAW assistant in Mulberry

Fig. 2 shows the flow behind the recipient suggestion tool (simplified for demonstration purpose).

The recipient suggestion RAW

Fig. 2. The "recipient suggestion" RAW

The flowchart shows where LLM inference actually sits: one step out of several and one the workflow may never reach. Two branches answer the question before it, and after it a confidence check decides whether the suggestion is shown at all. No model can change this flow, it is defined in the source code.

Conclusion

Retrieval-augmented workflows didn't come out of a whiteboard planning session. During the design phase for our assistive tools, we kept seeing the same pattern emerge, and realized existing terms didn't quite fit. Giving the pattern a name just made it easier to talk about and easier to say no to the things it isn't.

The practical payoff is control. A fixed execution graph can be read, tested, and audited like any other code path. A structured result carries the "why" behind its suggestion, so an operator can instantly see the reasoning and decide whether to trust it. Neither of those things survives when you hand the wheel over to the model.

None of this rules agents out. There are problems where the path to a solution genuinely can't be known in advance, and for those, autonomy is the right tool. But most of what our users need help with is not that. It is work they already know how to do, done faster, with the evidence laid out in front of them. For that, the workflow belongs to us and the decision belongs to them.