Blog — MLOps / Deployment

Gunicorn:
FastAPI in production.

Gunicorn is a widely used WSGI server for running Python web applications. Paired with Uvicorn workers, it's the standard way to get FastAPI production-ready — worker count, load balancing, and graceful shutdown.

Dwayo Team Jun 15, 2023
1. Configuration

Gunicorn configuration.

Gunicorn is a widely used WSGI server for running Python web applications. When deploying FastAPI, Gunicorn is often used in conjunction with Uvicorn to provide a production-ready server.

Install Gunicorn using:

pip install gunicorn

Run Gunicorn with Uvicorn workers:

gunicorn -k uvicorn.workers.UvicornWorker your_app:app -w 4 -b 0.0.0.0:8000

Here:

How many workers to run.

The -w flag in the Gunicorn command determines the number of worker processes. The optimal number depends on factors like CPU cores, available memory, and the nature of your application.

For example, on a machine with four CPU cores:

gunicorn -k uvicorn.workers.UvicornWorker your_app:app -w 4 -b 0.0.0.0:8000

If your application performs a significant amount of asynchronous I/O operations, you might increase the number of workers. However, keep in mind that too many workers can lead to resource contention.

3. Load balancing

Load balancing and scaling.

In a production setting, deploying multiple instances of your FastAPI application and distributing incoming requests across them is essential for scalability and fault tolerance. The number of worker processes can impact the optimal scaling strategy.

Consider using tools like nginx for load balancing or deploying your application in a container orchestration system like Kubernetes.

Shutting down without dropping requests.

Ensure that Gunicorn handles signals gracefully. FastAPI applications may have asynchronous tasks or background jobs that need to complete before shutting down. Gunicorn's --graceful-timeout option can be set to allow for graceful termination.

gunicorn -k uvicorn.workers.UvicornWorker your_app:app -w 4 -b 0.0.0.0:8000 --graceful-timeout 60

This allows Gunicorn to wait up to 60 seconds for workers to finish processing before shutting down.

In conclusion, the choice of Gunicorn and worker processes is a crucial aspect of deploying FastAPI applications in a production environment. Fine-tuning the number of workers and configuring Gunicorn parameters according to your application's characteristics and deployment environment ensures optimal performance and scalability.

Related

Where this fits.

Production deployment decisions like this are part of our MLOps work — getting the serving layer right before scale turns a small misconfiguration into an incident.

Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll help you think through production readiness for your AI or API serving layer.

Book a free 30-min call → More articles →