Untangle.bio
untangle.bio是一款面向生物技术领域的自助式B2B SaaS,可生成下游处理路线、模拟分离过程,并运行技术经济分析(CAPEX/OPEX/投资回收期),可通过网页应用或直接从Claude通过其MCP服务器访问。
文档
Documentation
untangle.bio is an AI-native platform for downstream process design in biotechnology. Generate optimal purification routes, run real-time simulations, and perform techno-economic analysis — all in one workspace.
Quick Start
Get up and running in under 5 minutes. Open the app — the canvas greets you with an empty state and a single call to action: Start here. Click it to launch the guided wizard, which walks you through feed definition, target selection, and route generation in one flow.
🚀 Click "Start here"
The canvas empty state shows a single button. Click it to open the guided wizard — no setup required.
🔬 Define Feed & Targets
Enter your feed components from 300+ molecules, set flow rate, and select the products you want to recover.
📊 Review & Pick a Route
Browse ranked routes by yield, purity, and CAPEX. Apply one to the canvas, inspect it, then return to pick another.
Pro tip: Start with the Balanced optimization mode for your first project. It provides a good mix of yield and purity while keeping costs reasonable.
Scope & Limitations
untangle.bio is a conceptual process design tool, intended for early-stage route screening and feasibility assessment — not for detailed engineering or final process validation. Understanding what the simulator does and does not model will help you interpret results correctly.
What the simulator does
- Steady-state mass balances across each unit operation based on separation efficiency, rejection coefficients, and component properties
- Flow and concentration tracking through every stream in the flowsheet, including branched flowsheets wired the way you drew them
- pH-dependent solubility checks and precipitation warnings
- Yield and purity estimates for each product at every step
- Indicative capital and operating cost ranges based on literature-derived correlations
- Dynamic fermentation kinetics in the bioreactor: growth, oxygen transfer and product formation over time decide the titer (see Bioreactor & Fermentation)
- In Thorough mode: temperature- and pressure-aware streams, heating and cooling duties with steam and cooling water demand, enforced component mass closure, and recycle loops converged by Wegstein iteration
What the simulator does not do
- No rigorous phase equilibria — Thorough mode adds Raoult-law bubble points with boiling point elevation and latent-heat duties, but activity coefficients and equation-of-state calculations are not performed. Real mixture non-idealities (salting-out effects, co-precipitation, ternary phase diagrams) are not captured.
- No detailed transport modelling — concentration polarisation, fouling kinetics, and gel-layer effects in membrane operations are approximated by fixed rejection parameters rather than solved from first principles.
- No chromatography band profiles — chromatographic separations are represented by an overall recovery and purity factor, not by breakthrough curves or plate-height models.
- No downstream reaction kinetics — fermentation is modelled kinetically in the bioreactor, but enzymatic conversion, degradation and aggregation during purification are not.
- Simplified crystallisation — the crystallised fraction is capped at the thermodynamic limit set by the supersaturation ratio at the crystalliser temperature, with a user-defined recovery fraction as the kinetic limit under it. Nucleation and growth kinetics, crystal size distribution and supersaturation profiles over time are not modelled.
- No hydrodynamics — pressure drops, pump sizing, pipe velocities, and fluid dynamics are outside scope.
Engineering interpretation required. Results should be treated as indicative order-of-magnitude estimates. Promising routes identified by untangle.bio should be validated with detailed process modelling, pilot-scale experiments, and consultation with separation specialists before making engineering or investment decisions.
Start Here Wizard
The recommended entry point for new sessions. Click Start here on the empty canvas (or Start here in the toolbar) to open the guided wizard. It bundles feed definition, target selection, and route generation into a single step-by-step dialog so you can go from a blank canvas to a ranked list of routes in minutes.
Wizard steps
- Feed stream — set flow rate, pH, temperature, and add components from the molecule database or manually.
- Target products — select one or more components you want to recover. For multiple products, the generator finds branching routes that separate each product.
- Generation settings — choose optimization goal (diversity, balanced, yield, purity, simplest, or high selectivity), and minimum yield and purity thresholds.
- Pick a route — results stream in live. Browse the list, inspect the 3D yield, purity and cost-of-goods scatter plot, and click Apply to canvas on the route you want to explore.
After applying a route: The canvas is populated with the full process flowsheet and you can inspect every stream and unit operation. The Back to Results button in the toolbar lets you return to the results list, remove the current route from the canvas, and pick a different one. See Back to Results for details.
Platform Workflow
untangle.bio follows a proven engineering workflow that mirrors how process engineers actually work — from initial feed characterization to final economic evaluation.
1. Feed Stream Definition
Define your input stream with volumetric flow rate and component specifications. The platform includes an extensive molecule database with physical properties for accurate modeling:
- Flow rate: L/hr, with automatic unit conversion
- Components: Concentration (g/L), molecular weight, charge state
- Properties: pKa, isoelectric point, diffusion coefficient
- pH and temperature: Critical for precipitation modeling
2. Target Product Selection
Select one or multiple target products from your feed components. untangle.bio optimizes routes for maximum recovery and purity of specified products, with support for complex multi-product separations.
3. Route Generation
The AI engine generates thousands of candidate routes using a diversity-preserving genetic algorithm. Set constraints and optimization goals:
- Minimum yield: Typically 10-90% depending on application
- Minimum purity: Product specification requirements
- Optimization mode: Yield, purity, balanced, or high selectivity
- Results: Each route arrives with yield, purity, CAPEX and cost per kg of pure product, so the list can be ranked on economics as well as on quality
4. Simulation & Mass Balance
Run rigorous mass balances with stream-level tracking of concentrations, pH, and flow rates. The simulation engine handles:
- Conservation of mass and volume at every node
- pH propagation through mixing and chemical addition
- Precipitation warnings based on solubility limits
- Real-time feasibility checking
Results panel: When a unit operation node is selected, the Properties panel opens with the Results tab active by default — showing stream concentrations, yield, purity, and flow rate at a glance. Switch to Parameters to adjust operating conditions.
5. Refine and Stress-Test
A picked route is where the engineering starts. Three tools work on the flowsheet as it stands on the canvas:
Back to Results
After applying a generated route to the canvas you may want to compare it visually with alternatives before committing. The Back to Results button in the toolbar makes this frictionless:
- Apply a route to the canvas from the wizard or generator dialog.
- Inspect the flowsheet — check stream labels, unit operation results, and the properties panel.
- Click Back to Results in the toolbar to remove the current route from the canvas and reopen the results list with all previously generated routes still shown.
- Pick a different route and apply it, or re-run generation with new settings.
The results list is saved in memory for the current session. It is cleared when you start a new project or close the browser tab.
Unsaved changes: Clicking Back to Results removes the currently applied route from the canvas. Any manual edits made after applying (added nodes, changed parameters) will be lost. Use Ctrl + Z after returning if you change your mind.
Building Your Own Process Flow
Alongside the automated route generators, you can build and edit process flowsheets entirely by hand — drag nodes onto the canvas, wire them together, and run the simulation yourself. This is useful when you want to test a specific sequence, reproduce a literature process, or make targeted modifications to a generated route.
Step 1 — Place a Feed Stream
Drag a Feed Stream node from the left palette onto the canvas. Double-click it to open the feed configuration dialog. Set the volumetric flow rate, temperature, and pH, then add your components — either from the built-in molecule database or as custom entries with manually entered properties.
Step 2 — Add Unit Operation Nodes
Drag one or more unit operation nodes from the palette. Available operations are grouped by category:
- Upstream / Fermentation: Stirred tank bioreactor, fed-batch bioreactor, perfusion bioreactor, continuous bioreactor (chemostat), air-lift fermentor, continuous heat sterilizer
- Reaction & Mixing: Mixing vessel
- Clarification: Disc centrifuge, depth filtration, microfiltration, flocculation/coagulation, Nutsche filter, basket centrifuge
- Purification: Ultrafiltration (10 kDa / 30 kDa MWCO), cation exchange, anion exchange, affinity, size exclusion, hydrophobic interaction (HIC), reverse osmosis, electrodialysis, distillation, liquid-liquid extraction, precipitation
- Polishing: Nanofiltration, reverse phase, crystallization, activated carbon adsorption, viral inactivation
- Drying: Spray drying, freeze drying, vacuum tray drying, thin-film evaporator, fluid bed dryer
A separate Sources & Sinks section at the top of the palette holds the feed, product and waste nodes plus reagent feeds — wash water (💧), NaOH solution (🔵), and HCl solution (🔴). These are covered in steps 1, 4 and 5.
Double-click any unit operation to configure its parameters (MWCO, pH target, wash volume, etc.).
Step 3 — Switch to Connect Mode
Press C to enter Connect Mode (the current mode is shown in the status bar at the bottom of the window). In this mode, hovering over a node reveals its connection handles. Click and drag from one handle to another to draw a stream edge.
| Handle | Position | Meaning |
|---|---|---|
output | Right side of feed node | Feed stream outlet — connect to the first unit operation's input |
input | Left side of unit operation | Main process inlet |
light | Right side of unit operation | Light-phase outlet — permeate, filtrate, mother liquor, volatiles |
heavy | Bottom of unit operation | Heavy-phase outlet — retentate, concentrate, solid, crystals |
dilution | Top of filtration nodes | Auxiliary water inlet for diafiltration — connect a Wash Water node here |
Tip: Press V to return to Select Mode for moving nodes around. Use Ctrl + Z / Y for undo/redo.
Step 4 — Add Product and Waste Nodes
Every outlet of every unit operation must terminate at either a Product node or a Waste node — the simulator validates this before running. Drag these from the palette (Sources & Sinks section) and connect them to the appropriate outlets.
- Connect the outlet carrying your target product to a Product node
- Connect all other outlets to a Waste (Wastewater Treatment) node
- Multiple unit operations can share the same Waste node
Step 5 — Add Reagent Feeds (optional)
To model diafiltration or pH adjustment, drag reagent feed nodes from the palette and connect them to the appropriate inlets:
- 💧 Wash Water → connect to the
dilutionhandle on any filtration node - 🔵 NaOH Solution → connect as an additional inlet for pH increase
- 🔴 HCl Solution → connect as an additional inlet for pH decrease
Step 6 — Run the Simulation
Press F5 or click the Recalculate button in the toolbar. The simulator performs a steady-state mass balance through every node in sequence, propagating concentrations, flow rates, and pH along every stream. Results appear as labels on stream edges and as summary panels on each unit operation node.
Validation errors: If the simulator reports dangling outlets or unconnected streams, check that every outlet handle on every unit operation is connected to either a downstream node, a Product node, or a Waste node. Unconnected outlets prevent the simulation from running.
Feed Definition
Accurate feed characterization is critical for reliable route optimization. untangle.bio provides comprehensive tools for defining complex biotechnology feeds.
Component Database
The platform includes 300+ pre-characterized molecules across key categories:
- Proteins: Antibodies, enzymes, therapeutic proteins
- Organic acids & amino acids: Citric, acetic, lactic, and others
- Sugars & polysaccharides: Glucose, sucrose, complex carbohydrates
- Salts & alcohols: Buffer components, ionic species, solvents
- Cells: E. coli, CHO, yeast with size distributions
- Small molecules: Vitamins, antibiotics, lipids, polyphenols, terpenes
Database integration: Clicking any molecule automatically populates all relevant properties for separation modeling, including molecular weight, charge, and transport properties.
Route Generation
untangle.bio uses advanced algorithms to explore the vast space of possible purification sequences and identify optimal routes based on your criteria. The evolutionary algorithm is the default and recommended mode — it returns only feasible routes (meeting both yield and purity thresholds) and streams results live to the UI as each simulation completes.
Genetic Algorithm Approach
The platform employs a diversity-preserving genetic algorithm optimized for breadth rather than convergence:
- Population size: 600 genomes for maximum diversity
- Generations: 15 for single-product runs, 40 for multi-product, each with fresh injection (~14% new genomes per generation)
- Selection: Tournament selection with low elitism (2%)
- Mutation: Multi-type operations (add, remove, replace steps)
- Only feasible routes returned: Routes failing yield or purity thresholds are excluded. If zero feasible routes are found, the top 5 infeasible routes are shown with clear warnings.
Economics on every route
Each candidate route is costed as it is generated, so the results list carries a screening-grade cost of goods in dollars per kg of pure product alongside yield, purity and CAPEX. The 3D scatter plot uses cost of goods as its third axis, switching to a log scale when the spread across routes is wide. These numbers are built on one shared set of default economics, which makes them comparable between routes rather than a quote for any one of them. Open the Economic Analysis dialog on the route you keep for the full picture.
Optimization Goals
Choose the optimization goal to guide the search:
- Diversity (default) — explore the widest range of process options with maximum variety
- Balanced — good mix of yield and purity (Yield × Purity)
- Yield — maximize mass recovery of product
- Purity — maximize product concentration relative to impurities
- Simplest — favor shorter routes with fewer unit operations
- High Selectivity — only allow routes where every step enriches the product over impurities (selectivity > 1 at every step)
Expert Rules Integration
All generated routes pass through 30+ expert rules that eliminate physically impossible or economically infeasible combinations:
- Chromatography requires water ≥ 500 g/L and particles < 1 g/L; prior clarification is required when cells are present
- Membrane operations require appropriate particle size reduction first
- Crystallization requires supersaturation (concentration > solubility limit)
- Size-based separation must follow large-to-small ordering
- Reverse phase blocked for proteins (MW > 1500 Da) — causes denaturation
- Organic acids must be at low pH (<6) for liquid-liquid extraction to work
- Nutsche filter and basket centrifuge require solids > 2–5 g/L
- Fluid bed dryer requires granular feed or prior solid-forming step
Simulation Engine
The simulation engine performs rigorous mass and energy balances with real-time validation of process feasibility and stream compatibility.
Mass Balance Methodology
untangle.bio uses a mass-flow-based approach for accurate modeling:
// Convert to mass flows
mass_flow = concentration × volumetric_flow
// Apply separation efficiency
retained_mass = mass_flow × rejection_coefficient
permeate_mass = mass_flow × (1 - rejection_coefficient)
// Enforce conservation
total_out = retained_mass + permeate_mass
assert(total_out == mass_flow_in)
pH Tracking
pH is tracked throughout the entire process with buffer capacity weighting:
- Volume-weighted mixing of streams
- Chemical addition effects (NaOH, HCl)
- Precipitation warnings near isoelectric points
- Henderson-Hasselbalch equation for acid solubility
Thorough Simulation
Calculate solves the flowsheet as a mass balance. Thorough, next to it in the toolbar, replays the same flowsheet with a heavier engine: streams carry temperature and pressure, thermal steps are costed from real duties, and recycle loops are converged rather than ignored. Nothing about your flowsheet changes, so the two runs are directly comparable.
What Thorough adds
- Thermodynamic streams: Raoult bubble points including boiling point elevation, so an evaporator or a dryer knows the temperature it actually runs at.
- Rigorous thermal duties: sensible and latent heat per step, converted into steam demand, cooling water and utility cost per hour.
- Enforced component mass closure: every component is checked in and out of every step, not just the total flow.
- Converged recycle loops: tear streams are solved by Wegstein iteration until the residual settles. See Recycle Loops.
- Equipment design basis: each step reports its sized duty (vessel volume, membrane area, column volume), how many parallel units that needs, and pump and agitation power with the pressure drop behind it.
Reading the results
The dialog opens with a convergence banner (iterations, method, residual), then plant totals: heating kW, cooling kW, electricity kW, steam kg/h, cooling water m³/h and utility cost per hour. Below that, one row per step shows inlet temperature, heating, cooling, electricity, steam and the mechanism that set the duty. Click a row to expand its design basis, every outlet stream with flow, temperature, pH, mass flow and composition, and any warnings raised for that step.
An unconverged run is not a solved flowsheet. If the banner says NOT CONVERGED, the numbers below it are a snapshot of an iteration that never settled. Reduce the recycle fractions or simplify the loop and run it again.
No baseline yet? Opening the dialog on a flowsheet that has not been calculated runs the baseline mass balance for you. If that fails, the dialog says what is blocking it (missing target product, unconnected outlet, unconfigured feed) instead of running on nothing.
Recycle Loops
Real plants send material backwards: mother liquor to the crystallizer, retentate to the feed tank, solvent to the extractor. Add a loop from the Thorough dialog with + Add recycle and give it four things:
- From step — the operation the stream is drawn from.
- Outlet — the light or the heavy outlet of that step.
- To step — any earlier step, or the same one.
- Fraction — how much of that outlet goes back, up to 0.99.
The loop is drawn on the canvas: a recycle mixer node is spliced in ahead of the target step and the return stream is wired back into it. After a run, the loop edge is labelled with the converged flow, temperature and pH, so you can see what is actually circulating. Loops live on the canvas rather than in the dialog, which means they survive closing the dialog and reloading the page.
The overlay is display only. Calculate still sees the same acyclic chain it always did, so adding a loop never invalidates your baseline mass balance.
Loops change the economics. Recycled material is not effluent. Once a thorough run has converged, the Economic Analysis panel prices waste water on the net load leaving the plant and shows thorough flows + duties (recycle loops priced) in its header. With a loop drawn but no converged run it says loops not priced, run Thorough, and treats every outlet as if it left the process.
Bioreactor & Fermentation
The bioreactor is a dynamic fermentation model, not a yield table. It integrates growth over time and reports the broth that purification will actually receive.
What is modelled
- Growth: a selectable rate expression — Monod, Haldane/Andrews substrate inhibition (the default, with its maximum at S = √(K s K i)), or Contois, whose half-saturation rises with biomass so a dense or flocculent culture limits itself — with oxygen limitation, maintenance on both substrate and oxygen, and cell death.
- Product formation: Luedeking-Piret, split from growth, so growth-coupled and non-coupled products behave differently. The titer is reported split between the two terms.
- Substrate distribution: Herbert-Pirt, q s = μ/Y XS + m s + q p /Y PS, reported as the fraction of every gram of sugar that went to growth, to upkeep and to product, with a closure check against the mass ledger. A long, slow, oxygen-starved batch spends a visibly larger share of its carbon on maintenance.
- Thermodynamic yield ceiling: the Gibbs energy dissipation correlation of Heijnen and van Dijken gives the maximum biomass yield the substrate allows from nothing but its carbon number and degree of reduction (236 kJ per C-mol biomass on glucose, giving 0.55 g DCW/g aerobically and 0.12 g/g fermentatively). A yield coefficient above that ceiling is flagged as impossible rather than used.
- Product location: extracellular, intracellular or inclusion bodies. An intracellular titer is bounded by what the biomass can carry (about 30% of dry cell weight by default) and travels with the cells through every clarification step until a disruption step releases it.
- Oxygen transfer: viscosity-dependent k L a with a dissolved oxygen controller capped by the installed motor power. Under-aerate the vessel and the titer falls, which is the point.
- Host organism: a library of production hosts, each with its own maximum growth rate, yields, saturation constants, maintenance, viscosity, cultivation temperature and pH window. Anaerobic yield pairs are held separately, so an ethanol fermentation is not costed on aerobic biomass yields.
- Your own host: pick Custom host and describe a strain outright — growth rate, yields on substrate and oxygen, saturation and inhibition constants, maintenance, death rate, oxygen demand, viscosity, pH window, cultivation temperature and cell size and density. Or start from a library organism and replace only the numbers you measured. Everything hand-entered is bounds-checked, the carbon balance is enforced, and the Gibbs-dissipation ceiling still applies — a biomass yield that thermodynamics forbids is flagged whether it came from the library or from you. Cell diameter and density carry downstream, where clarification is a Stokes and sigma-factor calculation on exactly those two numbers.
- Maintenance with temperature: the upkeep demand follows Tijhuis and Heijnen, m G = 4.5 exp[−(69000/R)(1/T − 1/298)] kJ per C-mol biomass per hour, which holds regardless of carbon source or electron acceptor. It roughly triples between 25 and 37 °C, so a warm fermentation loses yield with nothing else changed: the library coefficient is scaled from the temperature it was measured at, and the report shows both it and the value thermodynamics implies.
- pH on growth: μ max is scaled by a two-shoulder curve, f = 1/(1 + 10 pH_low−pH + 10 pH−pH_high), normalised so the organism’s own optimum is never penalised. The shoulders are half-rate points held per host, so A. niger keeps most of its rate at pH 2 where a CHO culture has none left. A pH setpoint is therefore a kinetic decision, not only a downstream one.
- Substrate: named explicitly from the molecule database rather than inferred. Methanol, acetate, glycerol or a hydrolysate all work, and residual substrate keeps its real identity downstream.
The titer ceiling
Product inhibition needs a P max, and where that number comes from is stated rather than buried. In order: the figure you type into Max Product Titer; for a colloidal product (a gum, a protein, a lipid) the top of its reported concentration range, because a gelling exopolysaccharide does not leave solution at a fraction of saturation; otherwise the lower of a fifth of aqueous solubility and the top of that range. That keeps butanol on its real ~15 g/L ABE ceiling while letting gellan gum reach the 10–20 g/L an industrial fermentation actually makes — the fifth-of-solubility rule alone pinned it at 2 g/L, from a listed solubility of 10 g/L. The ceiling drives growth continuously through the Levenspiel term, whose sharpness is also settable, so the batch slows towards it instead of being clipped at it, and the ceiling that bound a run is reported next to the titer.
Batch, fed-batch and continuous
- Batch — a charge is drawn down until it runs out, the growth stalls, or the harvest point arrives.
- Fed-batch — fed on a substrate trigger, at a constant rate, or exponentially to a growth-rate setpoint. The exponential feed is a substrate-stat: it never feeds faster than the culture is eating, which is what keeps residual sugar near zero instead of piling up the moment oxygen holds growth below the setpoint.
- Continuous — a real chemostat, integrated to steady state, where μ settles to D + k d and the residual substrate is set by the kinetics rather than by the feed. Above μ max the culture washes out and is reported as such; asked to, the model scans D and reports the dilution rate that maximises D × titer.
When the fermentation ends
The end of a batch is an economic decision, not only an exhaustion event, so the model makes it. It tracks cycle productivity as the batch runs — product in the vessel over the whole cycle it takes to get there, turnaround included — and harvests once that has passed its peak. The alternatives (run to substrate exhaustion, to maximum productivity, or to a fixed time) are selectable, and the reason the run ended is always named: substrate exhausted, oxygen starved, product inhibition, vessel full, productivity optimum, steady state, washout or time limit.
Heating and cooling duty
The jacket has to remove the agitator as well as the biology. Cooling duty is reported as metabolic heat plus shaft power, less the latent heat the exhaust gas carries away as it leaves saturated — and charged to the TEA on that basis, which it was not before: sizing the utility on metabolic heat alone understated it by the whole specific power, 2 to 5 kW/m³ on an aerobic microbial vessel. The two peaks coincide, because the DO controller ramps the agitator hardest exactly when oxygen demand is highest. A warning fires when the total exceeds the vessel’s installed cooling capacity, which is a design failure a jacket cannot argue with.
Scale-up parameters
Impeller speed is solved from the specific power asked for, not typed in, and everything that follows from it is reported: tip speed, impeller Reynolds and Froude numbers, the Nienow mixing time, the gas flow number against the flooding transition, k L a, the hydrostatic split between oxygen saturation at the surface and at the sparger, and the dissolved CO 2 a tall vessel accumulates. At constant P/V, tip speed rises with scale and mixing time rises faster; a scale translation names which quantity the chosen criterion cannot hold constant, because that is the one that will be blamed when the large-scale batch underperforms.
Two ways to set the titer
- Predicted — the kinetics decide biomass and titer from the substrate you feed and the yields you set. Used by the Start Here wizard, which has the substrate and nitrogen source pickers.
- Specified — you pin a measured titer and cell density, and the substrate demand is back-calculated. If the feed cannot supply it, both scale down and the run reports the feed concentration it would have needed. Used by the generator dialogs.
Inlet to purification
The generator dialogs show a live preview of the broth: a component table with concentrations and mass flows, product titer, cell density, batch time, cycle time, working volume and vessel count, and the mechanism that limited the batch (substrate, oxygen or inhibition). A fermentation profile chart plots biomass, substrate, product, growth rate and dissolved oxygen against time, so an oxygen floor is visible rather than inferred. Route generation searches downstream of exactly this stream.
Sizing is on the full cycle. Vessel volume is set by fill, sterilization, inoculation, fermentation, drain and CIP together, not by fermentation time alone. Turnaround is typically 20–40% of a microbial cycle, and ignoring it undersizes the plant.
The choice of host also propagates downstream. Clarification uses the organism’s own cell diameter and density in a Stokes and sigma-factor calculation, so a bacterial broth needs far more centrifuge capacity than a yeast or filamentous fungal broth to reach the same recovery. The difference shows up mostly in equipment count and cost rather than in yield, until the machine count is capped.
Techno-Economic Analysis
Built-in cost estimation provides immediate economic feedback on route alternatives using industry-standard methodologies.
Capital Cost (CAPEX)
Equipment costs are scaled using the power law, with an exponent that varies by operation type — generally following the "six-tenths rule" but calibrated individually to each technology class:
CAPEX = Base_Cost × (Flow_Rate / Reference_Rate)^n
Total_CAPEX = Σ(Equipment_Cost × Lang_Factor)
The exponent n is not a fixed 0.6 for all equipment — it is calibrated per technology class. Chromatography columns and membrane systems (area-limited equipment) scale more favourably than thermal or cryogenic systems. As a rough guide: membrane and column operations sit in the lower range (~0.55–0.65), mechanical separators in the middle, and drying operations — especially freeze drying — at the higher end (~0.70–0.75).
Lang factors (1.5–3.0×) account for installation, instrumentation, and auxiliary equipment based on operation complexity. All base costs are referenced at 100 L/hr feed rate (2026 USD).
Operating Cost (OPEX)
Annual OPEX is built up from several itemized components:
- Utilities: Electricity, steam, and cooling water — each with configurable unit prices
- Consumables: Resins, membranes, filters, and process chemicals
- Maintenance: 5% of FCI annually (industry standard)
- Labor: Base operators plus 0.25 additional operators per unit operation beyond two
- QC/QA: Facility-type-specific rate — pharma 50%, food & beverage 20%, chemical 15% of labor cost
Working Capital
Working capital is itemized rather than estimated as a flat fraction of FCI:
- Accounts receivable: 45 days of annual revenue
- Product inventory: 30 days of COGS
- Cash reserve: 2% of annual OPEX
TEA Analysis Views
Open the Economic Analysis dialog from the toolbar. Inside it, the analysis is split across several tabs:
- Summary / CAPEX / OPEX: Full detailed TEA — CAPEX and OPEX breakdowns with working capital and cost waterfall
- Scale-up: Cost vs. annual throughput curves — shows economies of scale
- Investor: NPV, IRR, and payback period under configurable revenue and discount rate assumptions
- Sensitivity: Tornado charts showing which cost drivers (flow rate, yield, price) have the largest impact on project economics
Optimize Parameters
A generated route is a starting point, not a finished design. The Optimize button searches the operating parameters of the flowsheet on the canvas for a better way to run it. No step is added, removed or reordered, so the answer is always the same process you were looking at.
Objectives
Pick what the search should optimize for:
- Lowest cost per kg — weighs capital, running cost, yield and purity in one number. The best default.
- Lowest running cost — consumables, utilities, labour, maintenance and effluent, for an annual budget rather than an up-front one.
- Lowest investment — total installed cost, for when the capital budget binds.
- Highest annual profit — revenue at the assumed selling price minus running cost. Unlike cost per kg, this rewards making more product.
- Highest purity and highest recovery — push one quality metric as far as this topology allows.
- Lowest energy — total shaft and thermal power, for when the energy bill or the carbon footprint binds.
- Least waste water — total effluent volume, for when a discharge permit is the limit.
- Reach a target cost — the goal seek, described below.
Guard rails
Four of the objectives are minimised perfectly by making no product at all, and the quality objectives are maximised by spending unlimited capital. The dialog therefore offers three floors, all on by default where they apply:
- Purity and recovery floors — hold quality at the baseline values, or at numbers you set, while cost is minimised.
- Economic guard — caps CAPEX and cost per kg at the baseline plus 10% headroom while a non-economic objective is searched. Without it, purity and energy objectives buy membrane area freely; runs measured before it existed inflated CAPEX by factors of 2.7 to 8.5 with no gain in the answer.
- Production floor — required when feed throughput is a search variable. Purity and yield are ratios, so they do not stop a search from shrinking the plant to a quarter of its size and calling the smaller bill an improvement.
Feed throughput as a variable
Optionally the search may move the feed rate between 0.25× and 4× the canvas value. Composition is never touched: that is your measured fact, and it is edited on the feed node. Throughput matters because raw materials dominate annual running cost, and with the rate pinned a running-cost search is partly minimising a constant.
Three possible outcomes
- Improved — a table shows every parameter the search moved, its baseline value and its new value. Apply writes them onto the canvas.
- Already optimal — the search converged and your settings were the best it found. Nothing to apply.
- Budget exhausted — the evaluation budget ran out before convergence. Raise the budget and run it again; this is not a statement that the flowsheet is optimal.
Single-product flowsheets only. Optimize and Tornado model a route as one linear list of steps with one outlet each. A multi-product flowsheet branches, so both buttons are disabled when more than one target product is selected.
Goal Seek
Choose Reach a target cost in the Optimize dialog and name a cost per kg. The answer comes in two stages, and they answer different questions:
- Stage 1 — settings. The search looks for operating parameters that reach your target. Throughput is opened up by default here, because scale is often the honest answer to a cost target. If a setting exists, you get it, and you can apply it.
- Stage 2 — requirements. If no setting reaches the target, the tool reports what would have to become true instead: the titer, resin capacity or fouling behaviour needed, and how far each is from today. Those are R&D targets, not dials.
A goal seek is not bound by the economic guard, since cost is the thing being targeted.
Tornado Analysis
The tornado tool answers “which input actually moves the number?” and it does so in two distinct modes, which can be run together against any of the eight objectives with a variation percentage you choose:
- Operating levers — the parameters you control, each varied around its current value. This tells you where tuning pays and where it does not.
- Assumption uncertainty — the model calibrations, economic assumptions and feed titer the answer rests on, each varied across its plausible range. This is the error bar on the estimate, not a to-do list.
Bars are sorted by swing, with the low and high case shown either side of the baseline, so the dominant driver is the top bar. Keeping the two modes apart matters: a lever you can turn and an assumption you merely believe should never sit in the same ranking.
A separate cost-driver tornado (flow rate, yield, selling price) lives on the Sensitivity tab of the Economic Analysis dialog, for project-level economics rather than flowsheet parameters.
Multi-Product Routes
untangle.bio supports complex separations where multiple valuable products are simultaneously recovered from a single feed stream through branching routes. Each product is tracked individually for yield and purity and exits at a dedicated product node.
Branching Logic
At every two-outlet unit operation, each product is assigned to whichever physical stream carries more of its mass — heavy (retentate/solid) or light (permeate/filtrate). Products that end up in different streams at the same step are considered separated at that step and branch into their own product nodes. Products that remain together continue downstream together.
- Heavy outlet: retentate, concentrate, solid, precipitate
- Light outlet: permeate, filtrate, mother liquor, volatiles
- Each product is tracked through every step individually
- A product node is created at the step and outlet where each product first separates
Metrics
- Yield — total mass recovery: mass of all target products recovered / mass of all target products in feed
- Purity — best individual product purity achieved, measured at each product's own exit step (excluding water)
Design constraint: For N selected products, the route must produce exactly N distinct product nodes — each product must exit through a unique (step, outlet) combination. Routes that fail to separate all products are automatically rejected.
Multi-Product Generation Algorithm
Multi-product routes are generated by the same diversity-preserving genetic algorithm used for single-product runs — the generator simply switches to a multi-product fitness function when you select two or more targets. There is no separate constructive search; the wizard and generator both call the evolutionary streaming engine for any number of products.
Multi-product fitness
When more than one target is selected, a route's reported yield and purity are the average across all target products. A route that recovers one product well but fails to separate a co-product therefore scores low rather than being hidden — it still appears in the results list and 3D plot, flagged as infeasible, so you can see why it fell short.
- Each product is tracked individually through every step of the simulated route.
- Branching is encoded in the genome's outlet-handle genes — at each two-outlet operation a product follows whichever stream (heavy or light) carries more of its mass, letting products split into their own product nodes.
- The algorithm runs a larger 40-generation search for multi-product problems (versus 15 for single-product) to explore more branching combinations, with relaxed default thresholds (min purity 60%, min yield 40%) to account for the added complexity.
Simulation & validation
Every candidate genome is passed through the full mass-balance simulation engine and gated the same way as single-product routes:
- Expert rules — evaluated against the actual stream composition at each step inlet (not just the feed). Violations are rejected immediately.
- Products must separate — routes are scored on how well each target ends up at a distinct outlet; those that keep products together score poorly on the averaged purity/yield.
- Purity & yield thresholds — routes are streamed to the UI as they finish, tagged as feasible or infeasible relative to your targets.
Streaming results: Routes are yielded to the UI as soon as each simulation completes — you see results appear live without waiting for the full population to finish.
Branched Flowsheets
A multi-product flowsheet is not a straight line: a centrifuge sludge goes one way and its supernatant another, and each branch has its own downstream train. The canvas sends every step together with the stream it is actually fed, meaning the upstream step and which of its outlet handles the edge leaves from.
What follows from that:
- Simulation reads each step’s inlet from the drawn edge, so a branch fed from a heavy outlet is not silently handed the previous step’s light stream.
- An outlet that a later step consumes is treated as an internal branch, not as a product leaving the process, so yields are not double counted.
- Expert rules are checked against each step’s true inlet and its own wired ancestor chain. A microfiltration fed centrifuge sludge is no longer rejected for “following” an ultrafiltration on an unrelated branch.
- Thorough runs take branch flows net of any recycle draw on the same outlet.
- The flow mismatch warning compares each step’s simulated inlet against its wired feeder, and names both.
Route generation and the parameter search still work on sequential routes. Wiring is honoured wherever a drawn flowsheet is simulated.
Expert Rules System
The platform incorporates decades of downstream processing knowledge through 30+ expert rules that prevent infeasible designs. Rules are evaluated against the actual stream composition at each step — not just the feed — so violations caused by upstream operations are also caught. The rules below are representative examples, not the full set.
Core Prerequisites
- Chromatography prerequisites: Requires water ≥ 500 g/L and particles < 1 g/L; prior clarification is required when cells are present
- Membrane fouling prevention: UF/NF with cells requires prior clarification
- Crystallization thermodynamics: Concentration must exceed compound-specific solubility (from molecule database)
- Drying constraints: Must be the final operation in any route
- Consecutive duplicate rejection: Same unit operation type cannot appear back-to-back
- Concentrate before drying: Solids must be > 50 g/L before any drying step
Practical Rules (additional examples)
- No reverse phase for proteins: Blocked for MW > 1500 Da — causes irreversible denaturation
- Maximum chromatography steps: No more than 3 chromatographic operations (cost and cycle time)
- Redundant clarification: Maximum 2 consecutive clarification operations
- Size-based separation order: Must follow large → small (centrifuge → MF → UF → NF)
- Concentrate before crystallization: Requires ≥ 50% of compound solubility limit
- LLE requires low pH for organic acids: At pH > 6, organic acids are fully ionized (A⁻) and won't partition into organic phase
- LLE not after drying: No aqueous phase remains after a drying step
- Nutsche filter requires solids: Needs > 5 g/L solids in feed
- Basket centrifuge requires solids: Needs > 2 g/L solids in feed
- Fluid bed dryer requires granular feed: Needs prior solid-forming step or > 50 g/L solids
Selectivity
Selectivity (α) measures how well each unit operation enriches the target product relative to impurities. It is calculated at every step and shown in the route results panel.
α = (product concentration factor) / (impurity concentration factor)
concentration factor = C_out / C_in
α > 1 → step enriches product over impurities (good)
α = 1 → no selective separation
α < 1 → step enriches impurities more than product
Use High Selectivity optimization mode to restrict results to routes where every single step achieves α > 1.0. The route list and 3D scatter plot include filter tabs to show only routes with fully monotone selectivity profiles.
Unit Operations Reference (38 total)
All 38 operations have 3–5 configurable parameters (accessible by double-clicking the node) with validated scientific defaults used by the process generator.
| Category | Operations | Outlets |
|---|---|---|
| Clarification | Disc centrifuge, depth filtration, microfiltration, flocculation, Nutsche filter, basket centrifuge | 2 (light + heavy) |
| Purification | UF 10k, UF 30k, cation/anion exchange, affinity, size exclusion, HIC, reverse phase, precipitation, distillation, reverse osmosis, electrodialysis, liquid-liquid extraction | 2 (light + heavy) |
| Polishing | Nanofiltration, crystallization, activated carbon adsorption, viral inactivation | 2 for NF/crystallization; 1 for viral inactivation, activated carbon |
| Drying | Spray drying, freeze drying, vacuum tray drying, thin-film evaporator, fluid bed dryer | 2 (solid/heavy + volatiles/light) |
| Cell disruption | High-pressure homogenizer, bead mill | 1 (single outlet) |
| Upstream / reaction | Stirred tank, fed-batch, perfusion, continuous (chemostat), air-lift bioreactors; conversion reactor; continuous heat sterilizer; mixing vessel | 1 (single outlet) |
Two-outlet operations produce a heavy stream (retentate, concentrate, solid, crystals) and a light stream (permeate, filtrate, mother liquor, volatiles). Both outlets must be connected to a downstream node, product node, or waste node before the simulation will run.
Molecule Database
The built-in molecule database currently covers a limited set of common biotech components — proteins, sugars, organic acids, amino acids, salts, alcohols, and cell types. It is actively being expanded over time based on user feedback and real-world process cases.
For testing purposes: If your molecule is not in the database yet, you can add it manually directly in the feed stream dialog. Enter the component name and as many physical properties as you know (MW, charge, solubility, pKa, log P, etc.). The simulator will use whatever properties you provide — missing values are handled gracefully, though accuracy improves with more complete data.
Note that molecules added this way are local to your simulation only — they are not automatically added to the central database. To request a molecule be added for all users, reach out via LinkedIn.
Property Categories
- Basic: MW, charge, typical concentrations
- Solubility: Water solubility, pH-dependent solubility
- Transport: Diffusion coefficients, viscosity effects
- Thermodynamic: Heat capacity, formation enthalpy
- Chemical: pKa values, log P, isoelectric points
Pulling in a molecule from PubChem
If a compound is not in the database, search it by name in the molecule picker and pull its properties straight from PubChem. Molecular weight, formula and the physical properties PubChem holds are filled in for you; anything it does not carry stays editable, and the app tells a rate-limited lookup apart from a compound that genuinely has no record.
Suggest a Molecule
The database is continuously expanding. If you work with a molecule that is missing, reach out on LinkedIn — feedback from practitioners directly shapes what gets added next.
Exports & Reports
Two export buttons sit at the right of the toolbar, both enabled once a simulation has results:
- CSV material balance — every stream in long format, a per-step summary, and the techno-economic figures if a TEA has been calculated. Made for pasting into a spreadsheet model rather than for reading.
- Printable report — opens a formatted process report in a new tab; save it as PDF from your browser’s print dialog. It covers the flowsheet, stream table, per-step results and, once a TEA has been calculated, the economics.
Pop-ups: the report opens in a new tab, so allow pop-ups for untangle.bio. If the browser blocks it, the output panel says so.
Calculate the TEA before exporting if you want economics included. Both exports take whatever the session holds at the moment you click.
Account & Projects
Signing in gives you an Account button in the toolbar and a dedicated account page.
- Cloud projects — Save and Open work against your account rather than only the browser, so a flowsheet follows you between machines. Saving under an existing name overwrites that project. File export stays available as an escape hatch and for sharing.
- Units and currency — set your preferred units and currency on the account page and the whole app displays results that way, including the dialogs and exports.
- Session persistence — the canvas survives navigating between pages and reloading the tab, so opening the docs or the molecule library does not cost you your work.
AI Connector (MCP)
untangle.bio ships a Model Context Protocol (MCP) server, so you can drive the engine — generate purification routes, simulate mass balances, and run techno-economic analysis — directly from your own AI assistant. Inference runs on your model and plan; the connector only answers tool calls, and never holds an API key on your behalf.
Endpoint: https://mcp.untangle.bio/mcp Copy (Streamable HTTP). The connection is authorized once via OAuth, after which the tools appear in your assistant's tool menu.
Using Claude
If you use Anthropic's Claude, the button below opens the Add custom connector dialog with the name and endpoint already filled in — just review and confirm, then complete the one-time OAuth sign-in. No copy-pasting the URL into settings.
On a Team or Enterprise plan? Individual members can't add custom connectors themselves — a workspace Owner must add it once for the organization first. The Owner uses this link: https://claude.ai/admin-settings/connectors?modal=add-custom-connector&connectorName=Untangle&connectorUrl=https%3A%2F%2Fmcp.untangle.bio%2Fmcp Copy After that, each teammate opens Settings → Connectors, finds Untangle (labeled "Custom"), and clicks Connect to authorize with their own account.
Using ChatGPT
ChatGPT doesn't yet support a one-click install link, so you add the server manually (a quick, one-time step). In ChatGPT, enable Settings → Connectors → Advanced → Developer mode, then Connectors → Add, give it a name, paste the endpoint https://mcp.untangle.bio/mcp, choose OAuth, and create. Via the Responses API, pass the same URL in the request's tools array as an mcp tool ({"type": "mcp", "server_url": "https://mcp.untangle.bio/mcp"}).
Manual setup (advanced & other MCP clients)
Because MCP is a shared, vendor-neutral standard, the same endpoint works from every MCP client — only the place you paste it differs. In Claude (claude.ai or the Claude Desktop app) you can also add it by hand: open Settings → Connectors → Add custom connector, give it a name, and paste the URL; Claude walks you through the OAuth sign-in on first use. For any other MCP client — Cursor, Cline, Zed, custom agents built on the MCP SDKs, or the reference MCP Inspector — register it wherever that client lists MCP servers, using the HTTP/SSE (Streamable HTTP) transport and the same URL. In every case no per-vendor build or API key is required on your side.
Available tools
| Tool | What it does |
|---|---|
list_unit_operations | List available downstream unit operations. |
get_molecules | List the built-in molecule database with physical properties. |
generate_processes | Evolutionary search for single-product purification routes. |
generate_processes_multiproduct | Branching flowsheets that recover 2+ products in parallel. |
simulate_separation | Step-by-step mass and energy balance for one route. |
simulate_fermentation | The bioreactor alone: the broth it makes, why it ended there, and charts of the run. |
calculate_tea | Detailed techno-economic analysis (CAPEX / OPEX / COGS / payback). |
tea_scale_analysis | Sweep economics across a range of throughput scales. |
tea_sensitivity | Which inputs move COGS most (tornado data). |
tea_investor_metrics | NPV, IRR and payback framed for an investor conversation. |
get_bioreactor_parameters | The settable parameters of one bioreactor, with units and ranges. |
flowsheet_link | Deep link that opens a route on the canvas, plus an importable JSON file. |
plot_process_landscape, plot_tea_sensitivity, plot_tea_scale, plot_cost_breakdown, plot_fermentation_profile | PNG charts of a prior result — route landscape, tornado, economics vs scale, cost breakdown, fermentation panels. |
report_issue | Your assistant flags a result that looks physically or economically implausible, straight to our engineers. |
Building your own agent? The complete machine-readable reference — every tool with its full JSON argument schema, the server instructions and the canonical workflow — is at untangle.bio/llms-full.txt. Integrating over plain HTTP instead of MCP? The curated OpenAPI spec of the engine endpoints is at untangle.bio/openapi.json (requests need a signed-in account's bearer token).
A typical flow: get_molecules → generate_processes → simulate_separation on a promising route → calculate_tea on the simulated result. With a fermentation upstream, start one step earlier: simulate_fermentation settles the broth first, since that broth is the feed every downstream route is conditioned on. It answers in charts — the same concentration, oxygen, volume and heat panels the workspace draws.
Under the Hood
The engine behind untangle.bio is documented as a knowledge graph: one node per topic — a separation model, a cost basis, a generator rule — cross-linked to the topics it depends on and to the code that implements it. The Untangle Universe renders that graph as an interactive 3D map: 91 knowledge nodes, 267 cross-links and 132 model files, clustered into constellations, with the most-connected nodes forming the visible backbone of the platform. A third view drops the topics for the molecule database itself: 5,467 stored property values across 333 molecules, one mote per value.
Ready to start designing processes? Launch the workspace and begin optimizing your downstream operations.