πΌοΈ Multimodal ProcessingΒΆ
This page explains how to use MMIRAGE with vision-language models (VLMs) to process datasets that include images.
Before reading this page, familiarise yourself with Concepts and the Quickstart text-only example.
OverviewΒΆ
MMIRAGE supports multimodal inputs natively. When a dataset sample includes an image reference, MMIRAGE resolves it to a PIL Image object and passes it alongside the text prompt to the VLM.
The key additions compared to a text-only pipeline are:
one or more
inputswithtype: imageimage_base_pathin the dataset config to resolve relative pathsa
chat_templateon the processor matching your VLMβs expected format
Dataset configuration for imagesΒΆ
Add image_base_path to the dataset entry in loading_params.datasets to
specify the directory where image files are stored:
loading_params:
datasets:
- path: /path/to/dataset
type: loadable
output_dir: /path/to/output/shards
image_base_path: /path/to/images
If your dataset already stores absolute paths in the image_path field,
you can omit image_base_path.
Image input variablesΒΆ
Declare image inputs in processing_params.inputs by setting type: image:
processing_params:
inputs:
- name: image
key: image_path
type: image
- name: question
key: question
The key extracts a value from each sample (via JMESPath).
For image inputs, this value is treated as a file path.
MMIRAGE resolves it as follows:
If
image_base_pathis set on the dataset config, the path is joined with that prefix.The resolved path is loaded as a PIL Image and stored under
name.
Inside your prompt template, you can reference the image variable by name. MMIRAGE places it in the correct position in the multimodal message:
outputs:
- name: answer
type: llm
output_type: plain
prompt: |
{{ image }}
Answer this question about the image:
{{ question }}
Chat templateΒΆ
VLMs expect their inputs in a specific format.
Set chat_template on the processor to enable the correct message structure:
processors:
- type: llm
server_args:
model_path: Qwen/Qwen2-VL-7B-Instruct
tp_size: 4
trust_remote_code: true
chat_template: qwen2-vl
default_sampling_params:
temperature: 0.1
top_p: 0.95
max_new_tokens: 768
The chat_template value is passed directly to the SGLang engine.
Common values:
Model family |
|
|---|---|
Qwen2-VL |
|
LLaVA-style |
|
InternVL |
|
Leave chat_template empty (or omit it) for text-only LLMs.
Complete multimodal exampleΒΆ
processors:
- type: llm
server_args:
model_path: Qwen/Qwen2-VL-7B-Instruct
tp_size: 4
trust_remote_code: true
chat_template: qwen2-vl
default_sampling_params:
temperature: 0.1
top_p: 0.95
max_new_tokens: 768
loading_params:
state_dir: /path/to/state
datasets:
- path: /path/to/image_dataset
type: loadable
output_dir: /path/to/output/shards
image_base_path: /path/to/images
num_shards: 8
shard_id: "$SLURM_ARRAY_TASK_ID"
batch_size: 8
processing_params:
inputs:
- name: image
key: image_path
type: image
- name: question
key: question
outputs:
- name: answer
type: llm
output_type: plain
prompt: |
{{ image }}
Answer this question about the image:
{{ question }}
output_schema:
question: "{{ question }}"
answer: "{{ answer }}"
image_path: "{{ image }}"
execution_params:
mode: slurm
account: my_account
job_name: vlm-pipeline
nodes: 1
ntasks_per_node: 1
gpus: 4
cpus_per_task: 32
time_limit: "05:59:59"
retry: true
merge: true
max_retries: 2
Batch size for image workloadsΒΆ
VLMs process significantly fewer samples per second than text-only LLMs. Typical guidance:
Start with
batch_size: 8and reduce if you hit OOM errors.Use
tp_sizematching the number of GPUs on your node.For large images, consider reducing
max_new_tokensto save memory.
See alsoΒΆ
Quickstart β text-only example to compare with
Pipeline β how image inputs flow through the mapper
Configuration Reference β full
loading_paramsandprocessorsreferenceSLURM & Cluster Deployment β running VLM jobs at scale