T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

EMNLP 2026 Submission

Shao-Jun Xia* Duke University
Huixin Zhang* Texas A&M University
Zhengzhong Tu Texas A&M University

* Equal contribution

Paper Code Google Drive Data Hugging Face Data

Motivation

Can visual in-context learning work when the demonstration and query come from different tasks?

Visual in-context learning (VICL) allows a model to infer an image transformation from input-output demonstrations without updating parameters. Existing VICL methods usually assume that the demonstration pair and the query belong to the same task. In practice, the available visual example can be mismatched: a user may provide a deraining example while the query requires dehazing, or a reflection-removal pair while the target is shadow removal.

This mismatch makes VICL ambiguous. The model must decide whether to imitate the demonstrated transformation or infer the target transformation from the query image. T2T-VICL is motivated by the observation that low-level vision tasks share latent relationships such as artifact suppression, contrast recovery, color adjustment, and illumination correction, and that these relationships can often be expressed more clearly in language than in raw pixels.

Illustration of specific-task VICL and cross-task VICL.
Illustration of Cross-Task Visual In-Context Learning and Its Relationship to Prior Work.

Prior work covers same-task visual-prompt VICL and same-task text-guided VICL, while task-transfer studies analyze task relationships without performing visual in-context generation. Cross-task VICL, where a Task-A demonstration guides a Task-B query, remains unexplored.

Main Contributions

T2T-VICL turns mismatched visual demonstrations into implicit task-difference prompts.

  1. Cross-task VICL formulation and benchmark. We formulate cross-task visual in-context learning for low-level vision and introduce a benchmark with implicit cross-task descriptions for studying how VLMs generalize across mismatched visual demonstrations.
  2. VLM -> sVLM -> VLM teacher-student pipeline. A large teacher VLM generates implicit task-relation descriptions, and a lightweight Qwen3-VL-4B student learns to generate content-dependent prompts for frozen image-editing VLMs.
  3. Score-based inference framework. Candidate outputs are selected with a hybrid evaluation strategy that combines attribution-enhanced prompting, visual instruction scoring, and IQA metrics.

Method

Framework: implicit relationship generation, knowledge transfer, and score-based inference.

T2T-VICL is a collaborative pipeline that leverages both large and small VLMs. A large Qwen2.5-VL-32B-Instruct teacher observes two paired tasks and produces structured implicit descriptions covering image content, visual changes, and task differences without explicitly naming the tasks. These descriptions are filtered for diversity and used as supervision for a Qwen3-VL-4B-Instruct student.

At inference time, the student only receives a Task A input-output demonstration and a Task B query. It generates an implicit prompt that guides a frozen image-editing VLM, such as Gemini-2.5-Flash, to perform the target transformation. The final output is selected from multiple candidates using VIEScore together with PSNR and SSIM.

Overview of the proposed T2T-VICL pipeline and workflow.
Overview of the Proposed Pipeline and its Workflow.

Dataset

12 low-level vision tasks across restoration, removal, and generation/enhancement.

The benchmark covers classic degradation problems, removal tasks, and generation/enhancement tasks. For each task pair, the teacher VLM produces implicit text relations; after semantic clustering and de-duplication, the dataset keeps 2,000 diverse descriptions per pair.

Category
Tasks
Restoration
Deblurring, Dehazing, Demoireing, Denoising, Deraining
Removal
Reflection Removal, Shadow Removal
Generation / Enhancement
Colorization, Harmonization, Inpainting, Light Enhancement, Style Transfer

Inference Example

The implicit prompt resolves what the fixed prompt leaves ambiguous.

Both the fixed-prompt and T2T-VICL settings receive the same visual inputs: a Task A input-output pair and a Task B query. The fixed prompt only states the structure of the task. T2T-VICL first asks the trained sVLM to produce a content-dependent implicit prompt, then passes that prompt and the images to the frozen VLM.

Detailed inference-stage example comparing fixed and implicit prompts.
Detailed Example of Cross-Task Visual In-Context Learning in Inference Stage.

Experiments

Top-tier cross-task results on Gemini 2.5 Flash and Seedream 4.0.

Table 3 compares twelve representative task pairs with a fixed-prompt baseline. Across these top-tier pairs, at least two of the three metrics consistently outperform the baseline for both models, while the remaining metric stays close. VIEScore improves throughout, highlighting the effectiveness of implicit prompt-driven generalization.

Top-tier quantitative results for Gemini 2.5 Flash and Seedream 4.0.
Average of evaluation metrics in the top-tier group, generated by Gemini and Seedream with Qwen prompts.

The qualitative examples below show how Gemini adapts across diverse cross-task pairs under the implicit prompt. The examples cover restoration-oriented pairs as well as generation/enhancement pairs, showing that the same prompting mechanism can transfer across perceptually related and more distant low-level tasks.

Representative Gemini examples for twelve top-tier cross-task pairs.
Representative Examples of Cross-Task In-Context Learning from Gemini in Twelve Pairs.

Second-Tier Results

VIEScore captures semantic consistency where PSNR and SSIM can be less decisive.

Across another nine task pairs, T2T-VICL consistently improves VIEScore over the fixed-prompt baseline on both Gemini and Seedream. PSNR or SSIM may decrease for several pairs, especially when the target task has one-to-many ambiguity such as colorization or style transfer. The paper therefore treats VIEScore as the primary signal for these second-tier cases.

Second-tier quantitative results for Gemini 2.5 Flash and Seedream 4.0.
Average of evaluation metrics in the second-tier group, generated by Gemini and Seedream with Qwen prompts.

The corresponding qualitative results remain visually consistent in most cases. These examples explain why a VIEScore gain can still indicate better task performance even when strict pixel-aligned metrics are slightly lower.

Representative Gemini examples for nine second-tier cross-task pairs.
Representative Examples of Cross-Task In-Context Learning from Gemini in Nine Pairs.

Additional Backbones

The prompt-transfer mechanism is evaluated beyond Gemini and Seedream.

The appendix evaluates open-source image-editing backbones including FLUX.2-dev, OmniGen 2, Qwen-Image-Edit, and FireRed. These models support multi-image inputs and edited-image outputs, making them compatible with the VICL setting. The results show that the T2T-VICL prompt mechanism is not tied to a single final generator.

Additional backbone results for FLUX.2-dev, OmniGen 2, Qwen-Image-Edit, and FireRed.
Average of evaluation metrics for additional backbones with fixed prompts and Qwen prompts.

Conclusion

T2T-VICL probes the boundaries of cross-task transfer in visual in-context learning.

T2T-VICL enables collaboration among multiple VLMs through an automatic reasoning mechanism. By converting heterogeneous visual demonstrations into implicit text prompts, the framework studies how VLMs jointly reason and adapt across tasks without task-specific fine-tuning of the final image-editing model.

Citation

BibTeX

@article{xia2025t2tvicl,
  title={T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs},
  author={Xia, Shao-Jun and Zhang, Huixin and Tu, Zhengzhong},
  journal={arXiv preprint arXiv:2511.16107},
  year={2025}
}