
This ComfyUI workflow leverages Google's Gemini model to demonstrate the capabilities of multimodal AI in generating coherent and contextually relevant text. The workflow utilizes several nodes, including the GeminiNode, which serves as the core processing unit, and the LoadImage node, which allows users to input images that the Gemini model can analyze and interpret. By integrating these components, the workflow showcases how Gemini's advanced reasoning capabilities can be applied to both text and image inputs, providing a seamless experience for generating text based on visual content. Additionally, the PreviewAny and BatchImagesNode nodes facilitate the visualization and management of multiple outputs, making it easier to handle and review generated content.
API
Use this workflow from code
Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.
GET https://comfy.org/workflows/download/f5feb0b58f78.jsonFetching workflow JSON…
Run it from Python or TypeScript with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.
# Install (beta)
pip install comfy-sdk # Python
npm i @comfyorg/sdk # TypeScript
# Run "Google Gemini" (Python)
from comfy_sdk import Comfy
client = Comfy(api_key="comfyui-...")
# This workflow, exported in API format (see note below)
wf = client.workflows.from_file("api_google_gemini_api.json")
asset = client.assets.from_file("input.png")
wf.set_input("2", "image", asset) # LoadImage
job = client.run(wf)
for output in job.get_outputs("20"): # SaveText
output.to_file(output.name)The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API). API access requires a Comfy Cloud plan with an API key. SDK docs · Get an API key
FAQ














