
This ComfyUI workflow leverages the powerful capabilities of the ElevenLabs model to transcribe speech from audio or video files into text. The workflow begins with the 'LoadAudio' or 'RecordAudio' node, which allows users to either upload an existing audio file or record new audio directly within the interface. Once the audio input is secured, the 'ElevenLabsSpeechToText' node processes the audio data, utilizing advanced speech recognition algorithms to convert spoken words into accurate text. This text is then made available for preview and editing through the 'PreviewAny' node, ensuring that users can review and refine the transcription as needed. This workflow is particularly useful for creating transcripts from interviews, lectures, or any spoken content, making it an invaluable tool for content creators, researchers, and professionals who need to convert speech to text efficiently.
API
Use this workflow from code
Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.
GET https://comfy.org/workflows/download/8d83904b8f5a.jsonFetching workflow JSON…
Run it from Python or TypeScript with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.
# Install (beta)
pip install comfy-sdk # Python
npm i @comfyorg/sdk # TypeScript
# Run this workflow (Python)
from comfy_sdk import Comfy
client = Comfy(api_key="comfyui-...")
wf = client.workflows.from_file("workflow_api.json")
job = client.run(wf)
for output in job.get_outputs("<output-node-id>"):
output.to_file(output.name)The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API). API access requires a Comfy Cloud plan with an API key. SDK docs · Get an API key
FAQ














