There is no dependable universal duration. “One use case” might be an existing model with clean data and a working endpoint, or it might require new data contracts, labelling, infrastructure, security approval, application changes, and months of outcome collection. Calling both projects a pilot does not make their timelines comparable.
Estimate the work in parts. First map data ingestion, validation, training, evaluation, artifact storage, deployment, monitoring, feedback, and ownership. Mark what already works and what must be built. Then identify external dependencies: access reviews, procurement, privacy assessment, labelled data, integration windows, and subject-matter experts. These often determine the schedule more than writing pipeline code.
A narrow pilot should prove one complete path rather than imitate an enterprise platform. Use one model, a defined dataset route, one target environment, explicit evaluation criteria, a controlled release, useful monitoring, and a recovery procedure. State what the pilot will not solve. If the team cannot yet measure real-world quality because labels arrive months later, say so and define the interim evidence.
Set milestones around observable results: a reproducible training run, a registered artifact linked to its evidence, a deployment in a test environment, an accepted production release, and a monitoring alert that reaches an owner. Add contingency for unknown data and integration problems. A date supplied before discovery should be treated as an assumption, not a commitment.
MLOps maturity continues after the pilot. More models, automated retraining, shared components, governance, and cost controls can be added as needs become clear. Ask for a range with assumptions and decision points rather than a confident “several weeks.” The honest schedule is the one derived from the current system, team availability, risk, and definition of done.