AI and software solutions
AI

Transcribing meeting recordings: how we built Transify's asynchronous job pipeline

A 500 MB meeting recording can't be processed in one HTTP request. How we designed queuing, job status, webhooks and signature checks in Transify.

Muhammed Göktuğ Temiz · 9 min read
Contents (12)
  1. The problem: a long job, a short request
  2. The job model: upload, get an ID, ask later
  3. The state machine: four states are enough
  4. The user chooses the scope of the job
  5. The stages of the pipeline
  6. Getting the result: polling or webhooks?
  7. Making webhooks safe
  8. Errors and limits are part of the contract
  9. How long are files kept?
  10. What is it built with?
  11. When does your own project need an asynchronous pipeline?
  12. Frequently asked questions

The video recording of a one-hour meeting can run to hundreds of megabytes; transcribing it, then translating and analysing it, takes minutes rather than seconds. Squeezing that into a single “send the file, wait for the answer” HTTP request makes it the most fragile part of the product. When we built our own product, Transify AI, we solved this with an asynchronous job pipeline. This article explains the design decisions, all of which you can see in Transify's public API, and the reasons behind them.

#The problem: a long job, a short request

Transify's API accepts audio or video files of up to 500 MB: MP3, WAV, M4A, OGG, MP4, MOV, MKV, AVI. Keeping a connection open while a file of that size is processed causes trouble in three places:

  • Timeouts: Browsers, proxies and load balancers cut connections that stay silent for too long. A connection dropped just before the end means starting over.
  • User behaviour: People close the tab, lock the phone, switch networks. The result must not depend on that connection.
  • Capacity: Queuing ten simultaneous uploads instead of processing them all at once keeps the load on the server and the AI model predictable.

The answer is a well-known pattern: accept the request immediately, hand back a job ID, do the work in the background and report the result through a separate channel.

#The job model: upload, get an ID, ask later

The upload request receives the file, queues the job and responds straight away. The response contains no result, only the job's ID and the address to ask for its status.

Request
curl -X POST https://transify.averissoft.com/api/v1/transcribe \
  -H "X-Api-Key: tk_live_..." \
  -F "file=@meeting.mp4" \
  -F "sourceLanguage=auto" \
  -F "targetLanguage=en" \
  -F "jobType=full" \
  -F "outputFormats=txt,pdf,srt"
Response (the message reads “Job queued”)
{
  "jobId": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "status": "queued",
  "statusUrl": "/api/v1/jobs/3fa85f64-...",
  "message": "İş kuyruğa alındı"
}

Using a random UUID as the ID is deliberate: sequential numbers (1041, 1042…) leak how many jobs have been processed and make it easy to guess someone else's. A job can also only be queried with the API key that started it; ask with a different key and the answer is simply “not found” rather than “this job isn't yours”, so even the job's existence isn't revealed.

#The state machine: four states are enough

StatusMeaningWhat the client should do
queuedThe job was accepted and is waiting its turnWait; don't upload the file again
processingThe pipeline is running; the response includes a progress percentageShow a progress bar if useful
completedOutputs are ready; download URLs are in the responseDownload the files
failedThe job could not be completedShow the error and retry if appropriate

Keeping the number of states small simplifies the client: anything other than completed or failed means “wait”. However many stages there are inside, the contract exposed to the outside stays the same, so adding a stage doesn't break anyone's integration.

#The user chooses the scope of the job

Not every recording needs translation or analysis. So the scope is set with the jobType parameter when the job is started:

jobTypeWhat is done
transcription_onlyTranscript only (default)
translationTranscript + translation into the target language
analysisTranscript + summary, decisions, action items
fullAll of the above

Separating the stages has two benefits: users don't wait for work they don't need, and each stage can be monitored and retried on its own. In Transify, speech recognition, translation and analysis are done by calling AI models through their APIs; the application's own job is to receive the file, run the steps in order and collect the result.

#The stages of the pipeline

  1. Acceptance and validation: File format and size are checked before any work starts; an unsupported format or an oversized file is rejected with a clear error code and never enters the queue.
  2. Speech to text: Whisper-based speech recognition. If no source language is given it is detected automatically.
  3. Translation (optional): The transcript is translated into the target language; more than 50 languages are supported.
  4. Analysis (optional): The summary, decisions, action items and key topics are extracted. More on this in extracting decisions and action items from meetings.
  5. Output generation: TXT, PDF and DOCX files, plus SRT/VTT for subtitles, are produced from the same content.
  6. Notification: The result is announced by email and, optionally for API users, by webhook.

#Getting the result: polling or webhooks?

There are two ways for a client to learn the result. Both are supported because each is right in different situations.

Status pollingWebhook
How it worksThe client asks for the job status at intervalsWhen the job finishes, Transify sends a POST request to your URL
When it fitsBrowser and mobile apps, quick experiments, systems not reachable from outsideServer-to-server integrations, large numbers of jobs
Weak pointGenerates unnecessary requests; with long intervals the result arrives lateNeeds a publicly reachable HTTPS URL and signature verification

#Making webhooks safe

A webhook is a request arriving at your server from outside; anyone who knows the URL could send a fake “job completed” notification. So four measures work together:

  • Signature: Every request carries an HMAC-SHA256 signature of the timestamp and the body in the X-Transify-Signature header. Only the two parties hold the signing secret.
  • Replay protection: Requests with a timestamp older than 5 minutes are rejected, so a captured notification can't be sent again later.
  • Retries: If your server doesn't return 2xx within 10 seconds, the notification is sent again after 30 seconds and after 5 minutes (three attempts in total). Failed jobs are reported too, with a job.failed event.
  • URL checks: The webhook URL must be HTTPS and resolve to a public IP. Addresses resolving to internal networks (127.x, 10.x, 192.168.x, 169.254.x) are rejected; otherwise the feature could be abused to make the server send requests into its own internal network (SSRF).
Python: verifying the webhook signature
import hmac, hashlib, time

def verify(body: bytes, signature: str, timestamp: str, secret: str) -> bool:
    # reject requests older than 5 minutes (replay protection)
    if abs(time.time() - int(timestamp)) > 300:
        return False
    expected = hmac.new(
        secret.encode(),
        f"{timestamp}.{body.decode()}".encode(),
        hashlib.sha256,
    ).hexdigest()
    return hmac.compare_digest(f"sha256={expected}", signature)

Comparing with hmac.compare_digest instead of == is not a detail: a constant-time comparison stops the signature from being guessed from response times.

#Errors and limits are part of the contract

A well-designed API makes failures as predictable as successes. In Transify every error has the same shape: a description for humans and a stable code for programs to act on.

Error response (“File exceeds the 500 MB limit”)
{ "error": "Dosya 500 MB sınırını aşıyor", "code": "FILE_TOO_LARGE" }
  • Identity: MISSING_API_KEY, INVALID_API_KEY: the key is missing, invalid or revoked.
  • Permission: INSUFFICIENT_SCOPE: the key is not allowed to use this endpoint.
  • Quota: RATE_LIMIT_EXCEEDED, PLAN_LIMIT_EXCEEDED, KEY_QUOTA_EXCEEDED: the hourly request limit, the plan limit or the key's own quota has been reached.
  • Input: FILE_TOO_LARGE, UNSUPPORTED_FORMAT, INVALID_WEBHOOK_URL: the request is rejected before it reaches the queue.

The hourly request limit applies per key. Its purpose is not to restrict users but to stop one runaway loop from affecting everyone.

#How long are files kept?

A meeting recording is sensitive data, and every file kept longer than necessary is an unnecessary risk. Files are deleted automatically after 30 days (720 hours) by default; on the Pro and Team plans you can choose deletion after 24 hours. Asking for an expired file returns FILE_NOT_FOUND. Where the data is processed, and the on-premises option, are covered in a separate article: on-premises AI and data protection.

#What is it built with?

The backend is written in .NET (ASP.NET Core), data lives in PostgreSQL and the system can be installed with Docker. Whisper is used for speech recognition and a MiniMax model for analysis, both reached through their APIs. A machine-readable definition of the API (OpenAPI) is published alongside the documentation and can be imported straight into Postman or code generators.

#When does your own project need an asynchronous pipeline?

This pattern isn't only for audio files. If any of the following applies, separating the work from the user's request is almost always the right call:

  • The operation takes longer than a few seconds: report generation, bulk email, image or video processing, large file imports.
  • The work depends on an external service whose speed you don't control: payments, shipping, an AI model.
  • Simultaneous requests can exceed capacity and need to be queued.
  • If the work is interrupted, it has to be safe to retry.

If you would like a similar set-up for your own processes, see our AI solutions and custom software and automation services, and our release and monitoring routine on the how we work page.

Frequently asked questions

What is an asynchronous job pipeline?

A structure in which long-running work is separated from the user's request and executed in the background: the request is accepted immediately, a job ID is returned and the result is collected later by polling or notification.

Which file formats and sizes does the Transify API accept?

MP3, MP4, WAV, M4A, OGG, MOV, MKV and AVI files of up to 500 MB each.

Should I use webhooks or status polling?

Webhooks are more efficient for server-to-server integrations; polling is more practical in browser and mobile apps or in systems that can't be reached from outside.

Who can use the Transify API?

API access is available on the Pro and Team plans; a key is created on the profile page and sent with every request in the X-Api-Key header.

  • asynchronous job pipeline
  • audio transcription api
  • whisper transcription pipeline
  • webhook signature verification
  • job queue architecture
  • meeting transcription
AuthorMuhammed Göktuğ Temiz

Muhammed Göktuğ Temiz is the founder of Averis Soft. He builds corporate websites, booking platforms, admin panels, e-commerce and mobile apps, running the work end to end from analysis and design through development, launch and server management. He works with Next.js, React, Node.js, Flutter and Docker; the Connect2Taxi booking platform for the Dutch market, İstanbul Sivasspor's club website with its admin panel and an appointment system for clinics are among his projects.

About usLinkedInGitHub
Blog

AI and more articles

All articles

Extracting summaries, decisions and action items from meetings: how we designed the AI analysis

A transcript alone doesn't get work done. How Transify extracts the summary, decisions, action items and owners as structured data.

On-premises AI and data protection: where should meeting recordings be processed?

Meeting recordings are personal data. Cloud versus on-premises AI, explained with the Transify example and a practical checklist.

Chatbots for small businesses: where they help and where they don't

Where AI assistants genuinely help small and medium-sized businesses, their limits, choosing a channel and a step-by-step set-up plan.
Contact

Let's bring your next project to life.

Tell us briefly about your project or book an online call; we'll get back to you within 24 hours.

WhatsApp