Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond

  • Full time
  • Meredale, Gauteng, South Africa View on Map
  • posted 23 hours ago
  • Posted : August 6, 2026 -Accepting applications

Job Description

A Visual Language Model (VLM) learns a joint understanding of image and text and generates text based on this understanding. Yet when multiple visual cues within an image conflict—such as a written word and its ink color—we do not fully understand how the model decides which signal to prioritize. A classical psychological paradigm to study how conflicting cues affect decision is the Stroop test, where participants are shown words in incongruent ink colors (e.g., the word “red” written in blue) and are instructed to report the ink color rather than read the word. We adapt the Stroop paradigm to VLMs and study how conflicting cues in the written word or ink color influence model behavior. Applying the Stroop test on a range of contrastive and generative VLMs suggests the models favor textual cues over color when text and color conflict. Analyzing the representation of the two cue types suggests that text cues in images are more salient than the color cues. This difference in saliency also translates to different intervention success in steering the VLMs: we found that it is easier to steer the embedding to make the model favor text cues than color cues. Overall, using the Stroop test, our findings suggest VLMs, similar to humans, are biased to “read” an image rather than to “see,” and the saliency of the two cue types is reflected in their embedding space. We will release our dataset and code to support future research upon acceptance.

Show more

Show less

Required skills

Related Jobs