What are the basics of model deployment?
What happens during model deployment
A trained model is a file of learned weights plus the code needed to run it. Deployment moves that model out of a notebook and into an environment where other software can send it data and get predictions back. The most common pattern is a web API that accepts requests and returns results.
Before launch, you usually fix the model version, document its input format and test it with realistic data. Dependencies such as exact Python package versions should be pinned so the same code behaves the same way on every server.
- Save the trained model in a portable format
- Wrap it in a small API service
- Package the service and its dependencies in a container
- Test latency and output with sample requests
- Roll out the new version gradually
Common serving options
Many teams start with a simple REST service built with a Python framework such as FastAPI. This works well for moderate traffic. For heavier loads, dedicated model servers such as TensorFlow Serving, TorchServe or NVIDIA Triton can handle batching and GPU use more efficiently.
Cloud providers also offer managed endpoints that handle scaling and infrastructure for you. Those services are convenient but can cost more at high volume, so compare pricing on the provider's current page before choosing.
Monitoring after deployment
A deployed model can slowly get worse as real-world data changes. This is often called data drift. Track request volume, error rates, response times and a sample of prediction quality over time.
Keep the previous model version available so you can roll back quickly if accuracy or latency drops after an update.
Common mistakes
- Deploying a model without pinning its library versions, which can change results between servers.
- Skipping monitoring because the model passed tests before launch.
