For retrieval you need overlapping text chunks with a pointer back to the source and time. Every transcription tool here can return chunks directly.
outputs: ["chunks"] with chunkSize in characters.from apify_client import ApifyClient # pip install apify-client
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("fguiraud/audio-video-transcriber").call(run_input={
"sources": [
{
"url": "https://archive.org/download/gettysburg_johng_librivox/gettysburg_address_64kb.mp3"
}
],
"outputs": [
"chunks"
],
"chunkSize": 1000
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
for chunk in item.get("chunks", []):
print(chunk)
No Python? Open the Actor page, paste your input in the form and press Start; results export to JSON, CSV or Excel.
Fields from a real run on 2026-09-30 (long values shortened):
{
"source": "https://github.com/openai/whisper/raw/main/tests/jfk.flac",
"status": "ok",
"language": "en",
"durationSeconds": 11,
"text": "And so my fellow Americans ask not what your country can do for you ask what you can do for your country"
}
Same as the transcription itself: $0.006 per minute (files), $0.003 per YouTube transcript.
Which vector database?
Any: the output is plain JSON you embed with your own model.
Can an AI agent use it?
Yes. It is an MCP server at https://mcp.apify.com/?tools=fguiraud/audio-video-transcriber (listed in the official MCP registry), and there is an agent skill for Claude Code, Codex and Cursor. Field-name guesses such as url or urls are accepted.