Vision-Language Models
August 26, 2026
GGSS: Steering Bias Away Without Retraining Generative Vision–Language Models
Yiqun Sun, Junyu Chen, Pengfei Wei, Lawrence B. Hsieh
arXiv

Generative vision–language models (VLMs) can describe images, answer visual questions, and support decisions that combine perception with language. But in human-centered settings, they can also produce systematically different answers when the only controlled change in an image is a demographic attribute such as perceived race or gender.

Retraining a large model every time a new bias is identified is expensive and often impractical. Inference-time debiasing offers a more flexible alternative: keep the model frozen and adjust its internal representations only when it is used.

The difficulty is that most existing inference-time methods were designed for static embeddings or CLIP-like models. A modern generative VLM does not represent an image with one vector. It passes many visual tokens into a language model, and those tokens carry different kinds and amounts of information. Applying the same hard correction to every token can remove useful visual content, distort activation geometry, or even degrade generation.

Our new paper introduces GGSS—Geodesic-Gated Spherical Steering, a lightweight inference-time framework designed specifically for this setting. GGSS learns a multi-dimensional bias subspace from counterfactual images, identifies which visual tokens carry unusually strong demographic signal, and gently rotates those tokens toward a debiased direction while preserving their original norms. The base VLM remains frozen throughout.

‍

Why generative VLMs need a different debiasing primitive

Adapting earlier debiasing methods to generative VLMs exposes three challenges.

First, activation geometry matters. Simple Euclidean subtraction changes both the direction and the length of a token vector. Downstream attention and decoding layers were trained on the model’s original activation distribution, so those changes can disrupt general capability—especially at larger model scales.

Second, bias is not distributed uniformly across visual tokens. Tokens associated with a face or clothing may carry demographic information, while tokens representing the background, pose, or text in the image may carry little or none. A uniform projection treats all of them as equally relevant and risks over-correcting useful content.

Third, many protected attributes are not one-dimensional. Binary gender may sometimes be approximated by one direction, but perceived race is multi-class. A practical method therefore needs to model a low-dimensional subspace rather than assume a single universal bias axis.

GGSS addresses these three problems with spherical geometry, token-level gating, and counterfactual subspace discovery.

‍

Introducing GGSS

GGSS has an offline discovery stage and a lightweight inference-time stage.

1. Discover a counterfactual bias subspace

In the offline stage, we pass sets of carefully matched counterfactual images through a frozen VLM. Within each set, demographic attributes vary while identity, occupation, pose, clothing, background, and lighting are kept as fixed as possible.

We pool and normalize the resulting visual activations, map their differences into a tangent space on the unit hypersphere, and apply singular value decomposition. The leading components form a multi-dimensional subspace that captures demographic variation in the model’s representation.

This approach learns directly from counterfactual activation shifts. It does not require training a new classifier or modifying the VLM’s parameters.

2. Measure bias signal token by token

At inference time, GGSS examines each visual token independently. It measures how strongly the token projects onto the discovered bias subspace and compares that magnitude with statistics collected during discovery.

This produces a calibrated gate for every token. Tokens with unusually strong protected-attribute signal receive more steering; tokens with little such signal remain close to their original representations. The intervention is therefore selective rather than uniform.

3. Steer along the sphere

For each token, GGSS constructs a target direction with the learned protected-attribute coordinates removed. Instead of jumping directly to that target with a hard Euclidean update, it uses spherical linear interpolation (Slerp) to travel along a geodesic arc.

After the rotation, GGSS restores the token’s original radius exactly. In other words, the method changes direction while preserving norm. A single strength parameter, α, controls how aggressively the model moves toward the debiased target.

The core idea is simple: steer only where needed, and follow the geometry the model already uses.

Overview of the GGSS framework.

‍

Evaluation

We evaluate GGSS on four generative VLMs spanning three model families and parameter scales from 4B to 12B:

  • Pixtral-12B
  • LLaVA-1.6-Vicuna-7B
  • LLaVA-1.6-Mistral-7B
  • Qwen3-VL-4B-Instruct

The discovery stage uses 480 real, pixel-aligned counterfactual face images from the REFLECT/FOCUS dataset. Evaluation covers three complementary bias protocols: a race-conditioned salary/education multiple-choice task, a race-paired income comparison task, and a gender-gap test on nurse-versus-doctor classification.

We compare GGSS with ten adapted inference-time steering baselines from four method families—INLP, MeanDiff, BendVLM, and LEACE—as well as prompt-based mitigation. Every method intervenes at the same late vision-to-language projection layer and is evaluated with a single selected steering strength per model. We separately measure general multimodal capability on MMStar, a 1,500-question benchmark spanning perception, reasoning, mathematics, and science.

‍

Results: lower bias without sacrificing general capability

Across the full four-model comparison, GGSS delivers the strongest mean reduction and the strongest worst-case result. It is the most effective method on three backbones; a LEACE variant achieves the largest reduction on Pixtral but is less consistent on the other models. At GGSS's selected operating points, average bias changes relative to the unsteered model are:

  • ‍−55% on Pixtral-12B‍
  • −90% on LLaVA-1.6-Vicuna-7B‍
  • −80% on LLaVA-1.6-Mistral-7B‍
  • −60% on Qwen3-VL-4B

On individual tasks, GGSS reduces the Nurse/Doctor gender gap by up to 96%, the race-conditioned multiple-choice disparity by up to 84%, and race bias in pairwise comparison by up to 61%.

Just as importantly, MMStar accuracy stays within ±0.6 percentage points of the unsteered baseline across all race- and gender-steering evaluations. The bias reductions are statistically significant on three of the four model backbones, while the MMStar changes are statistically indistinguishable from the unsteered model in all eight steering runs.

The ablations help explain why. On Qwen3-VL-4B, adding the token gate improves over ungated spherical projection, and combining the gate with Slerp produces the strongest result. At 12B scale, the geometric choice becomes especially important: matched Euclidean variants significantly reduce MMStar accuracy, while spherical steering preserves it.

‍

A controllable intervention, not blanket erasure

A lower bias score is not useful if it comes from making the model unable to perceive demographic information at all. GGSS provides an explicit control for this trade-off through α.

At a moderate setting on Qwen3-VL, for example, race recognition remains within four percentage points of the baseline while the model already realizes 55% of the available reduction on the race-conditioned multiple-choice task. Steering one attribute also leaves recognition of the other tested attribute intact across all four backbones.

This controllability matters in deployment. Some applications should reduce the influence of demographic cues; others may legitimately need to describe an attribute. In the latter case, operators can use a moderate strength or turn steering off.

‍

What this means for practical VLM debiasing

GGSS suggests that inference-time debiasing for generative VLMs should not be treated as a global vector-removal problem. The representation is multi-token, the relevant signal is unevenly distributed, and the geometry of the intervention affects whether model capability survives.

The method is attractive for real systems because it:

  • works with frozen model checkpoints and requires no full-model retraining;
  • supports multi-dimensional protected-attribute subspaces;
  • calibrates steering strength independently for each visual token;
  • preserves token norms by construction; and
  • can be reconfigured as the target bias or deployment requirement changes.

At the same time, reduced benchmark bias is not the same as complete fairness. GGSS does not remove biased knowledge from model parameters, and our evaluation does not cover every architecture, language, domain, protected attribute, or distribution shift. The demographic categories used by the datasets are operational labels based on perceived attributes; they do not represent self-identification or the full diversity of human identity.

Any consequential deployment should therefore combine steering with broader auditing, explicit operating-point selection, human oversight, and monitoring for residual or newly introduced harms.

‍

Looking ahead

GGSS shows that adaptive geodesic steering can be a practical primitive for debiasing generative VLMs at inference time. Future work includes tuning-free strength selection, intersectional and multilingual evaluation, broader open-ended generation tests, and extensions to non-demographic bias subspaces.

More broadly, the results point to a useful design principle for model steering: when internal representations live on structured manifolds and different tokens carry different signals, interventions should respect both the geometry and the locality of the information being changed.

Code: github.com/dukesun99/GGSS

Keywords:
Vision–Language Models, AI Fairness, Inference-Time Steering, Multimodal AI,Representation Geometry