All posts

Scaling Django Backends: From Query Optimization to Reliable Background Tasks

A practical guide to designing robust Django backends by aligning ORM queries, Django REST Framework viewsets, and atomic background workers.

3 min read

When architecting a Django backend, performance issues rarely stem from Python being too slow. Most bottlenecks occur at the seams: leaky abstractions in the ORM, chatty Django REST Framework (DRF) serializers, or pushing long-running workflows directly into the request-response cycle. Structuring these layers cleanly from day one prevents performance degradation as your database grows.

Pruning Queries in the ORM Layer

Django's ORM is expressive, but it invites N+1 queries if you access related models within serializers or loops without explicit prefetching. To build clean interfaces, push database constraints and optimizations as close to the models as possible.

Instead of scattering filtering logic across views, encapsulate querying behavior inside custom QuerySet managers. This keeps query construction testable and DRY.

from django.db import models

class InvoiceQuerySet(models.QuerySet):
    def pending_delivery(self):
        return self.filter(
            status=Invoice.Status.PENDING
        ).select_related("customer").prefetch_related("items__product")

class Invoice(models.Model):
    class Status(models.TextChoices):
        PENDING = "PENDING", "Pending"
        PAID = "PAID", "Paid"
        CANCELLED = "CANCELLED", "Cancelled"

    customer = models.ForeignKey("Customer", on_delete=models.CASCADE, related_name="invoices")
    status = models.CharField(max_length=20, choices=Status.choices, default=Status.PENDING)
    total_amount = models.DecimalField(max_digits=10, decimal_places=2)
    created_at = models.DateTimeField(auto_now_add=True)

    objects = InvoiceQuerySet.as_manager()

Using select_related handles single-valued relationships via SQL joins, while prefetch_related avoids queries across many-to-many or reverse foreign keys. Combining both into dedicated querysets stops serializer performance leaks before they reach your controllers.

Designing Efficient DRF Endpoints

DRF makes exposing data straightforward, but standard nested serializers can undermine ORM optimizations if you fail to bind the queryset correctly. When rendering list endpoints, avoid instantiating nested serializers that trigger queries per record.

from rest_framework import serializers, viewsets

class InvoiceItemSerializer(serializers.ModelSerializer):
    product_name = serializers.CharField(source="product.name", read_only=True)

    class Meta:
        from .models import InvoiceItem
        model = InvoiceItem
        fields = ["id", "product_name", "quantity", "unit_price"]

class InvoiceSerializer(serializers.ModelSerializer):
    customer_email = serializers.EmailField(source="customer.email", read_only=True)
    items = InvoiceItemSerializer(many=True, read_only=True)

    class Meta:
        from .models import Invoice
        model = Invoice
        fields = ["id", "status", "total_amount", "customer_email", "items"]

class InvoiceViewSet(viewsets.ReadOnlyModelViewSet):
    serializer_class = InvoiceSerializer

    def get_queryset(self):
        return Invoice.objects.pending_delivery()

By overriding get_queryset to use the optimized manager method, serialization requires only two SQL statements regardless of whether the endpoint returns 10 invoices or 1,000.

Delegating to Background Workers Safely

HTTP requests must resolve swiftly. Actions like sending notification emails, generating PDF summaries, or syncing with third-party billing providers should run asynchronously through a background queue like Celery or Dramatiq.

A common pitfall is dispatching a background job inside a database transaction before the record commits. If the worker picks up the job before the database commits the transaction, it queries an ID that does not yet exist. Always use transaction.on_commit to avoid race conditions.

from django.db import transaction
from rest_framework import status
from rest_framework.decorators import action
from rest_framework.response import Response
from .tasks import process_invoice_pdf

class InvoiceViewSet(viewsets.ModelViewSet):
    queryset = Invoice.objects.all()
    serializer_class = InvoiceSerializer

    @action(detail=True, methods=["post"])
    def generate_pdf(self, request, pk=None):
        invoice = self.get_object()
        
        # Schedule the task only after the database transaction successfully commits
        transaction.on_commit(lambda: process_invoice_pdf.delay(invoice.id))
        
        return Response(
            {"status": "PDF generation queued"},
            status=status.HTTP_202_ACCEPTED
        )

Writing Resilient Tasks

In the task implementation, pass primitive arguments like integers or strings rather than serialized model instances. Stale state passed as serialized model instances leads to subtle data corruption when concurrent modifications occur.

from celery import shared_task
from django.core.exceptions import ObjectDoesNotExist
from .models import Invoice

@shared_task(bind=True, max_retries=3, default_retry_delay=60)
def process_invoice_pdf(self, invoice_id):
    try:
        invoice = Invoice.objects.select_related("customer").get(id=invoice_id)
    except ObjectDoesNotExist as exc:
        # Retry in case of replication lag or transient database issues
        raise self.retry(exc=exc)

    # Generate document and store the result
    invoice.render_pdf_to_storage()

This decoupled pattern keeps views responsive, limits memory pressure during serialization, and isolates operational failures to retryable background queues without breaking the user experience.

Have a project in mind?

I build and run web apps end to end.

Get in touch