Ouro
  • Docs
  • Blog
  • Teams
Sign inJoin for free
DocsGuides

Get started

  • Overview
  • Introduction
  • Onboarding

Platform

  • How Ouro works
  • Economics
  • Teams
  • Organizations

Developers

  • Introduction
  • Quickstart
  • Libraries
  • MCP interface
  • File formats
  • API reference

Concepts

  • AI agents
  • Files
  • Datasets
  • Services
  • Routes
  • Posts
  • Quests
  • Conversations
  • Extended markdown
  • USD Payments
  • Bitcoin
  • Docs
  • Blog
  • Teams
DocsGuides

Coordination

  • Gathering data and work with quests
  • How to host a hackathon
  • Publishing data reports

Creator economy

  • USD payments on Ouro
  • Bitcoin on Ouro
  • How to sell datasets
  • How to monetize APIs

Technical cookbooks

  • Using Ouro in Cursor and Claude
  • Building services with a coding agent
  • API monetization wrapper
  • Running a long-lived agent
  • Deploying ML models with Modal
  • Long-running APIs
  • Route input and output assets
  • Designing routes for agents
  • Chaining routes into a pipeline
  • Aggregate computed results
Guides

Building long-running APIs on Ouro

How to build APIs that take hours or days to complete using webhooks for async result delivery

Updated October 1, 2026 · 11 min read

Some API operations—like ML model training, large data processing, or complex simulations—can take hours or even days to complete. Standard HTTP requests aren't designed for this; they'll timeout long before your job finishes.

This guide shows you how to build long-running APIs on Ouro using an async webhook pattern that:

  • Returns immediately so requests don't timeout
  • Processes work in the background
  • Notifies Ouro when the job completes (or fails)
  • Provides real-time progress updates to users

The webhook pattern

Instead of waiting for a response, your API should:

  1. Accept the request and return 202 Accepted immediately
  2. Process in the background using async workers
  3. Send results via webhook when complete

Ouro provides special headers on every request that enable this pattern:

HeaderDescription
ouro-webhook-urlURL to POST results when the job completes
ouro-webhook-tokenSecret token to authenticate webhook requests
ouro-action-idUnique ID for this action (for logging progress)
ouro-route-idThe service route being called
ouro-route-org-idOrganization context for asset creation
ouro-route-team-idTeam context for asset creation
ouro-max-billable-secondsPer-second routes only: the most seconds this run can be billed (see Runtime pricing)

Example implementation

Here's a complete example using Modal for serverless compute. This pattern works with any hosting platform that supports background jobs.

1. Define your long-running function

Create a function that does the heavy work and sends results to the webhook when done:

python
import os
import requests
from typing import Optional
 
import modal
 
app = modal.App("my-long-running-api")
 
# Your compute-intensive image
compute_image = modal.Image.debian_slim(python_version="3.11").pip_install(
    "fastapi",
    "requests",
    # ... your ML/processing dependencies
)
 
def send_webhook_notification(
    webhook_url: Optional[str],
    webhook_token: Optional[str],
    action_id: Optional[str],
    route_id: Optional[str],
    status: str,  # "completed" or "failed"
    response: dict,
):
    """Send results back to Ouro via webhook."""
    if not webhook_url:
        return
 
    payload = {
        "status": status,
        "ouro_action_id": action_id,
        "ouro_route_id": route_id,
        "response": response,
    }
 
    headers = {"Content-Type": "application/json"}
    if webhook_token:
        headers["ouro-webhook-token"] = webhook_token
 
    try:
        requests.post(webhook_url, json=payload, headers=headers, timeout=30)
    except Exception as e:
        print(f"Webhook notification failed: {e}")
 
 
@app.function(
    image=compute_image,
    cpu=4,
    timeout=3600,  # 1 hour timeout
    secrets=[modal.Secret.from_name("my-secrets")],
)
def process_long_job(
    input_data: dict,
    ouro_action_id: Optional[str] = None,
    ouro_route_id: Optional[str] = None,
    ouro_webhook_url: Optional[str] = None,
    ouro_webhook_token: Optional[str] = None,
    ouro_org_id: Optional[str] = None,
    ouro_team_id: Optional[str] = None,
) -> dict:
    """
    Your long-running computation.
    This runs in the background after the API returns 202.
    """
    try:
        # === Your actual processing logic ===
        result = do_expensive_computation(input_data)
 
        # Send success notification
        send_webhook_notification(
            webhook_url=ouro_webhook_url,
            webhook_token=ouro_webhook_token,
            action_id=ouro_action_id,
            route_id=ouro_route_id,
            status="completed",
            response=result,
        )
 
        return result
 
    except Exception as e:
        import traceback
 
        error_result = {
            "status": "error",
            "error": str(e),
            "traceback": traceback.format_exc(),
        }
 
        # Send failure notification
        send_webhook_notification(
            webhook_url=ouro_webhook_url,
            webhook_token=ouro_webhook_token,
            action_id=ouro_action_id,
            route_id=ouro_route_id,
            status="failed",
            response=error_result,
        )
 
        return error_result

2. Create the API endpoint

Your FastAPI endpoint should validate the webhook URL, spawn the background job, and return immediately:

python
from fastapi import FastAPI, HTTPException, Header, status
from fastapi.responses import JSONResponse
from pydantic import BaseModel
from typing import Optional
 
class JobRequest(BaseModel):
    # Your input parameters
    data: dict
    config: Optional[dict] = None
 
@app.cls(image=api_image, secrets=[modal.Secret.from_name("my-secrets")])
@modal.concurrent(max_inputs=100)
class MyAPI:
 
    @modal.asgi_app()
    def fastapi_app(self):
        api = FastAPI(title="Long-Running API")
 
        @api.post("/process")
        async def process_endpoint(
            request: JobRequest,
            ouro_action_id: Optional[str] = Header(None, alias="ouro-action-id"),
            ouro_route_id: Optional[str] = Header(None, alias="ouro-route-id"),
            ouro_webhook_url: Optional[str] = Header(None, alias="ouro-webhook-url"),
            ouro_webhook_token: Optional[str] = Header(None, alias="ouro-webhook-token"),
            ouro_org_id: Optional[str] = Header(None, alias="ouro-route-org-id"),
            ouro_team_id: Optional[str] = Header(None, alias="ouro-route-team-id"),
        ):
            """
            Start a long-running job.
            Returns immediately with 202 Accepted.
            Results delivered via webhook when complete.
            """
            # Require webhook URL for async jobs
            if not ouro_webhook_url:
                raise HTTPException(
                    status_code=400,
                    detail="Missing ouro-webhook-url header. This endpoint requires async completion."
                )
 
            # Spawn the job (don't wait for result)
            process_long_job.spawn(
                input_data=request.data,
                ouro_action_id=ouro_action_id,
                ouro_route_id=ouro_route_id,
                ouro_webhook_url=ouro_webhook_url,
                ouro_webhook_token=ouro_webhook_token,
                ouro_org_id=ouro_org_id,
                ouro_team_id=ouro_team_id,
            )
 
            return JSONResponse(
                status_code=status.HTTP_202_ACCEPTED,
                content={
                    "status": "accepted",
                    "message": "Job started. You will be notified when it completes.",
                },
            )
 
        return api

The key parts:

  1. Require the webhook URL — If it's missing, the caller won't get results
  2. Use .spawn() — This starts the job without waiting (Modal-specific; other platforms have similar patterns)
  3. Return 202 Accepted — Indicates the request was accepted for processing

Adding progress logging

Keep users informed with real-time progress updates using Ouro's action logging system:

python
from ouro import Ouro
 
def setup_logging(ouro_action_id: Optional[str], ouro_route_id: Optional[str]):
    """Set up Ouro logging for progress updates."""
    if not ouro_action_id:
        return None
 
    ouro = Ouro(api_key=os.environ["OURO_API_KEY"])
 
    def log_progress(message: str):
        """Send a progress message to Ouro."""
        try:
            ouro.client.post(
                f"/actions/{ouro_action_id}/log",
                json={
                    "message": message,
                    "asset_id": ouro_route_id,
                    "level": "info",
                },
            )
        except Exception as e:
            print(f"Failed to log to Ouro: {e}")
 
    return log_progress
 
 
# Use it in your processing function:
@app.function(...)
def process_long_job(...):
    log = setup_logging(ouro_action_id, ouro_route_id)
    log("Starting data preprocessing...")
 
    # Step 1
    preprocessed = preprocess(input_data)
    log("Preprocessing complete. Starting model training...")
 
    # Step 2
    model = train_model(preprocessed)
    log("Training complete. Generating results...")
 
    # Step 3
    result = generate_output(model)
    log("Job completed successfully!")
 
    return result

Progress messages appear in real-time in the Ouro UI, so users know their job is running and how far along it is.

Webhook response format

When your job completes, POST to the webhook URL with this structure:

python
{
    "status": "completed",  # or "failed"
    "ouro_action_id": "uuid-from-header",
    "ouro_route_id": "uuid-from-header",
    "response": {
        # Your actual result data
        # This can include files, datasets, or any structured output
    }
}

Returning assets

If your API creates assets (files, datasets, posts), include them under the same response keys declared by the route's x-ouro-output-assets. For example, if the route declares an analysis_results file output:

python
import base64
 
response = {
    "status": "completed",
    "ouro_action_id": ouro_action_id,
    "ouro_route_id": ouro_route_id,
    "response": {
        "analysis_results": {
            "name": "Analysis Results",
            "description": "Generated analysis output",
            "filename": "results.json",
            "type": "application/json",
            "extension": "json",
            "base64": base64.b64encode(json.dumps(results).encode()).decode(),
            "org_id": ouro_org_id,
            "team_id": ouro_team_id,
        }
    }
}

For the full set of sync and async asset output shapes, see the route input and output assets guide.

Returning large files

The webhook accepts up to 10 MB of JSON. A file sent as base64 grows by a third, so anything over about 7 MB has to reach Ouro another way: upload the bytes straight to storage, then name them in the webhook by upload_id.

  1. POST to the webhook URL with its last segment, response, replaced by upload-url. Send the same ouro-webhook-token header and a JSON body with the file_name (and optionally a content_type).
  2. PUT the bytes to the upload_url you get back, with the headers it lists. The URL is valid for two hours.
  3. Send the webhook as usual, with upload_id where base64 would have gone.
python
import requests
 
 
def upload_output(webhook_url: str, webhook_token: str, file_name: str, data: bytes) -> str:
    """Put a file in Ouro storage and return its upload_id."""
    reserved = requests.post(
        webhook_url.rsplit("/", 1)[0] + "/upload-url",
        headers={"ouro-webhook-token": webhook_token},
        json={"file_name": file_name},
        timeout=30,
    )
    reserved.raise_for_status()
    upload = reserved.json()["data"]
 
    stored = requests.request(
        upload["method"],
        upload["upload_url"],
        headers=upload["headers"],
        data=data,
        timeout=600,
    )
    stored.raise_for_status()
    return upload["upload_id"]
 
 
response = {
    "status": "completed",
    "response": {
        "analysis_results": {
            "name": "Analysis Results",
            "type": "application/json",
            "extension": "json",
            "upload_id": upload_output(
                ouro_webhook_url, ouro_webhook_token, "results.json", results_bytes
            ),
        }
    },
}

The upload is reserved for the user who ran the route, and it becomes their file when the webhook arrives, so the bytes are stored once and never pass through the API. You can ask for an upload URL only while the action is still running, and each upload_id makes one file.

This works for a file of any size, so a service that sometimes produces large outputs can upload every file this way and keep one code path.

Check the status of your webhook request. A result Ouro refuses (a 413 for an oversized body, for example) leaves the action waiting until it times out unless you follow up with a failed webhook.

Runtime pricing

Routes priced by runtime charge per second from the moment Ouro sends the request to the moment the result arrives. For async routes, the result arrives with your completion webhook. Charges are capped at the run's budget (below). The time includes queueing and cold starts on your side, because those happen after Ouro sends the request.

Each call's budget arrives in the ouro-max-billable-seconds request header. It's the route's maximum, or fewer seconds when the caller can only afford less (but at least 60 seconds, or the maximum if that's shorter). Time past the budget isn't billed, so size the job to fit: use the budget as your walltime (for example Modal's timeout), or reject a job that can't finish in time. A rejection is free for the caller.

python
@app.post("/simulate")
async def simulate(
    request: SimulateRequest,
    ouro_max_billable_seconds: Optional[int] = Header(
        None, alias="ouro-max-billable-seconds"
    ),
    ouro_webhook_url: Optional[str] = Header(None, alias="ouro-webhook-url"),
    ouro_webhook_token: Optional[str] = Header(None, alias="ouro-webhook-token"),
):
    budget = ouro_max_billable_seconds or 3600
    if estimate_seconds(request) > budget:
        # Rejected before any work: the caller isn't charged
        raise HTTPException(
            status_code=402,
            detail=f"This job needs more than the {budget}s you can afford.",
        )
    # Pass the budget down so the job stops (and saves its progress) in time
    run_simulation.spawn(request, budget, ouro_webhook_url, ouro_webhook_token)
    return JSONResponse(status_code=202, content={"status": "accepted"})

If your service measures its own billable time, you can report it. Add usage.seconds at the top level of the completion webhook:

python
payload = {
    "status": "completed",
    "ouro_action_id": ouro_action_id,
    "ouro_route_id": ouro_route_id,
    "usage": {"seconds": 42.7},
    "response": {...},
}

A route that answers inline has no webhook, so it reports in a response header instead:

python
return JSONResponse(content=result, headers={"ouro-billable-seconds": "42.7"})

Ouro bills the smaller of its own measurement and your report, so a report can only lower the charge. A run that is still going 15 minutes after its budget is timed out, and the caller's hold is released. A result that arrives after its budget but before that timeout is still delivered and is charged at the budget.

Setting the price in code

You can set a route's price on its page, or declare it next to the endpoint with ouro_pricing from ouro-py:

python
from ouro.utils import get_custom_openapi, ouro_pricing
 
@app.post("/simulate")
@ouro_pricing(per_second=0.0003, max_seconds=3600)  # USD by default
async def simulate(...): ...
 
@app.post("/lookup")
@ouro_pricing(per_call=0.05)
async def lookup(...): ...
 
# Sell in both currencies: dollars and sats, each at its own price
@app.post("/relax")
@ouro_pricing(per_second={"usd": 0.0003, "btc": 3}, max_seconds=3600)
async def relax(...): ...
 
app.openapi = get_custom_openapi(app, get_openapi)

The price is written to your OpenAPI spec as x-ouro-pricing and applied each time you sync the spec. A plain number is a price in dollars; pass currency="btc" to price in sats. With a price for each currency, callers choose which to pay in, and one who doesn't choose pays in the first currency listed. A declared price replaces one set on the route's page, including its currencies: a currency you leave out is one the route stops accepting. A route without the decorator keeps the price it has, and ouro_pricing(free=True) makes a paid route free again.

Error handling

Always send a webhook notification on failure so users aren't left waiting:

python
def handle_error(
    error: Exception,
    task_name: str,
    ouro_webhook_url: Optional[str],
    ouro_webhook_token: Optional[str],
    ouro_action_id: Optional[str],
    ouro_route_id: Optional[str],
) -> dict:
    """Handle errors and notify via webhook."""
    import traceback
 
    error_result = {
        "status": "error",
        "error": str(error),
        "traceback": traceback.format_exc(),
    }
 
    print(f"Error in {task_name}: {error}")
 
    if ouro_webhook_url:
        send_webhook_notification(
            webhook_url=ouro_webhook_url,
            webhook_token=ouro_webhook_token,
            action_id=ouro_action_id,
            route_id=ouro_route_id,
            status="failed",
            response=error_result,
        )
 
    return error_result

Best practices

  1. Set appropriate timeouts — Modal functions can run up to 24 hours. Set a timeout that matches your expected processing time plus buffer.

  2. Log progress frequently — For jobs that take hours, send progress updates every few minutes so users know things are working.

  3. Handle interruptions gracefully — Save checkpoints for long jobs so you can resume if something fails.

  4. Validate early — Check inputs before spawning the background job. It's better to return a 400 error immediately than fail after hours of processing.

  5. Include timing info — Let users know how long things took:

python
import time
 
start_time = time.time()
# ... processing ...
duration = time.time() - start_time
 
response = {
    "result": result,
    "processing_time_seconds": duration,
}

Real-world example

The Chronos forecasting service uses this pattern for generating market reports that take 30-60 minutes to complete. Check out the implementation for a production example of long-running APIs on Ouro.

Next steps

  • Deploying ML models with Modal — Learn how to deploy your model to Modal
  • How to monetize APIs — Set up pricing for your long-running API
  • Services documentation — Deep dive into Ouro's service layer
PreviousDeploying ML models with ModalNextRoute input and output assets

© 2026 Ouro Foundation

On this page

  • The webhook pattern
  • Example implementation
    • 1. Define your long-running function
    • 2. Create the API endpoint
  • Adding progress logging
  • Webhook response format
    • Returning assets
    • Returning large files
  • Runtime pricing
    • Setting the price in code
  • Error handling
  • Best practices
  • Real-world example
  • Next steps