Skip to content
In the field

Where a persistent Python desk earns its keep

These are evidence-led concepts, not deployments we claim as our own. Each story borrows a published operating problem, then shows where PySmith sits: the always-on Python desk behind MCP tool calls — never the system that signs the decision.

The execution layer for AI agents that need persistent Python state.

Evidence-led conceptDecision support

Decision-grade finance

Keep evidence, computation and review history together long enough to inspect, challenge and reproduce.

5.7% → 3.7%
Top-20 retrieval failure with contextual embeddings; + BM25 2.9% (−49%); + rerank 1.9% (−67%). Lab, not fintech production.[1] Anthropic 2024
>98%
Advisor teams using the internal AI Assistant. Corpus 100k documents; reported access 20% → 80%. Published collaboration figures.[2] OpenAI / Morgan Stanley
>5,000
Bankers on Rogo; up to 10 hours/week saved; >50 million documents. Vendor-reported research and diligence, not automatic underwriting.[3] OpenAI / Rogo
Read the story →
Evidence-led conceptMulti-agent operations

Shift-scale mining

A persistent, inspectable Python environment for testing coordinated decisions across a full operational horizon.

+5.56%
603,840 vs 572,017 tons in a calibrated 12-hour / 50-truck simulation. Simulation only — not a production uplift.[1] Zhang et al. 2020
Mining-Gym
Python discrete-event testbed for dispatch RL. Research, not certified control.[2] Banerjee et al. 2025
No published %
BHP × Microsoft at Escondida: operator-facing recommendations. The announcement does not publish a measured improvement.[3] BHP 2023
Read the story →
Evidence-led conceptPlanning continuity

Living supply plans

Keep the planning computation available through every replan — people and solvers keep the decision.

>98%
On-shelf availability in Unilever’s Walmart Mexico CPFR pilot; >13 billion computations/day. Company-reported, no control group.[1] Unilever 2024
~40%
Less weekly analysis time in JD.com’s planning assistant; +22% plans within 5% accuracy; +2–3% fulfilment. Author-reported.[2] Qi et al. 2025
~93%
OptiGuide benchmark accuracy. The paper warns generated code can run and still be wrong.[3] Li et al. 2023
Read the story →

Agents route tool calls through MCP. PySmith is the always-on Python desk those calls land on — packages warm, state intact, bill predictable.

In the field

Agents call tools. Python lands on a desk.

If one of these operating loops is yours, request preview access. Bring the duty cycle, not a wish-list of autonomous miracles.