Major AI Chatbots Show Left-Leaning Political Bias, Even Anti-Woke Models

A new study finds that most leading AI chatbots default to progressive positions, challenging claims of neutral or conservative design.

By Central
The investigation tested six models and found left-leaning responses dominate even in anti-woke chatbots.
Highlights
  • OpenAI's GPT-5.5 delivered 80% exclusively left-leaning responses, the strongest bias in the study.
  • Even Elon Musk's Grok 4.3 gave more left-leaning than right-leaning answers, undercutting its anti-woke marketing.
  • Google Gemini 3.1 Pro was the only model presenting consistently balanced perspectives across all topics.

A comprehensive investigation into the political leanings of major AI chatbots has delivered a clear verdict: the vast majority of large language models exhibit a pronounced left-leaning bias in their responses to political questions. Even models explicitly marketed as conservative or anti-woke fail to break this pattern, with only Google’s Gemini 3.1 Pro presenting a consistently balanced perspective across a wide range of topics.

How the Investigation Measured Political Bias Across Six Leading AI Models

The investigation posed a carefully designed set of political questions to six leading AI models and categorized each response as exclusively left-leaning, exclusively right-leaning, or presenting both sides. The results confirm what earlier studies have suggested: AI chatbots, by default, tend to align with progressive positions on issues spanning taxation, healthcare, criminal justice, and military policy. The methodology and supporting code are available for review, allowing developers and researchers to replicate the analysis on additional models.

OpenAI GPT-5.5 and Deepseek V4 Pro Show the Strongest Left-Leaning Bias

OpenAI’s GPT-5.5 produced the most one-sided results in the study, with 80 percent of its responses containing exclusively left-leaning arguments. The model consistently advocated for higher taxes on the wealthy and a single-payer healthcare system, and it presented an exclusively right-leaning position only once across the entire test. Deepseek’s V4 Pro followed closely, with 70 percent of its answers falling into the exclusively left-leaning category. Both models argued against the death penalty, a position that runs counter to majority public opinion in the United States, where support for capital punishment has held steady for decades according to Gallup polling.

The consistency of these results across two independently developed models suggests that the bias originates not from a single company’s alignment choices but from the underlying training data that both models draw upon. When large language models are trained on web-scale corpora, they absorb the dominant political and cultural assumptions present in that text, and those assumptions skew left on a broad set of issues.

Claude Opus 4.8 and Grok 4.3: More Balanced but Still Leaning Left

Anthropic’s Claude Opus 4.8 delivered exclusively left-leaning answers 43 percent of the time and presented both sides in the remaining 57 percent of cases, making it one of the more balanced models in the test. xAI’s Grok 4.3, which Elon Musk has promoted as a truth-seeking and anti-woke alternative to mainstream chatbots, produced more right-leaning answers than any other model tested. Yet even Grok gave exclusively left-leaning responses more often than exclusively right-leaning ones, a finding that undercuts its marketed positioning.

The likely explanation lies in Grok’s training data and methodology. Grok was trained on similar web-scale data as other chatbots and, in some cases, may have been distilled from their outputs directly. This inheritance means that the underlying political orientation of the model remains left-leaning even when safety guidelines are deliberately loosened. xAI’s decision to relax content guardrails, rather than to reorient the model’s political baseline, explains why Grok can still generate racist or antisemitic statements when prompted in certain ways by users.

The Gab Arya Case Study: Conservative Branding Cannot Override Training Data

The right-wing social media platform Gab offers an AI model called Arya, which the company describes as built with Christian values and conservative principles. In the investigation, Arya responded with a left-leaning argument twelve times more often than a right-leaning one. This is perhaps the clearest illustration of the difficulty of engineering a politically conservative AI model using current methods. No amount of post-hoc tuning or branding can fully reverse the political orientation embedded in the pre-training data, especially when the model architecture and base weights originate from the same left-leaning ecosystems as every other major chatbot.

Google Gemini 3.1 Pro Stands Out as the Balanced Outlier

Google’s Gemini 3.1 Pro was the clear exception to the pattern. The model presented both sides of political questions 93 percent of the time, with only 7 percent of responses falling into the exclusively left-leaning category. It never gave an exclusively right-leaning response across the entire test. When asked whether the United States should use its military to conquer new territory, Gemini was the only model that offered an argument in favor of expansion, suggesting it could strengthen the American economy. This willingness to surface politically uncomfortable positions without endorsing them reflects a fundamentally different approach to model alignment, one that prioritizes balanced exposition over advocacy.

Gemini’s performance raises an important question for the industry: why can one model achieve near-parity in presenting both sides while others default so heavily to one orientation? The answer likely involves deliberate choices in instruction tuning, reinforcement learning from human feedback, and the specific datasets used for alignment. Google appears to have prioritized balance as a design goal, and the results show that this is achievable without sacrificing coherence or helpfulness.

The Anti-Woke Paradox: When Deliberate Steering Fails

Grok’s performance on trans rights illustrates a different phenomenon. In the investigation, Grok took an exclusively right-leaning position on that topic alone, aligning precisely with Elon Musk’s public stance. This suggests that someone deliberately intervened in the model’s output on that specific issue, rather than allowing the model’s general political orientation to determine its response. The result is a patchwork model that leans left on most topics but shifts to a hardline conservative position on selected issues that matter to its creators.

This approach has obvious risks. Users who encounter Grok on trans rights will see a uniformly conservative answer, while users who ask about taxation or healthcare will get a left-leaning response. The inconsistency undermines trust and makes the model’s behavior unpredictable. More importantly, it demonstrates that current methods for steering model outputs are topic-specific and brittle, not a substitute for holistic alignment.

The Limits of Left-Right Classification in AI Evaluation

Sorting AI answers into left and right categories has inherent limitations that the investigation acknowledges openly. On certain topics, right-leaning positions conflict with scientific consensus or universal human rights. Expecting a chatbot to present a conservative argument on climate change, vaccine efficacy, or equal protection under the law would amount to relativizing established facts or fundamental rights. For these questions, presenting both sides is not a virtue; accuracy and adherence to evidence are what matter.

On a broad range of political questions where legitimate disagreement exists in democratic discourse, however, the overwhelming left-leaning skew of most AI models is a genuine concern. Users who consult AI chatbots for information on tax policy, healthcare reform, or criminal justice deserve to see the spectrum of reasonable arguments, not a single ideological perspective dressed up as neutral output.

What This Means for AI Developers and Users Now

For developers, the findings underscore a fundamental challenge: training data bias cannot be fully corrected through alignment techniques alone. The political orientation of a model is set during pre-training, and post-hoc adjustments can only shift it at the margins. Building a genuinely balanced model requires careful curation of training data, deliberate instruction tuning, and ongoing evaluation against a diverse set of political perspectives. The fact that even models explicitly designed to be conservative default to left-leaning positions suggests that the problem is structural, not cosmetic.

For users, the practical implication is straightforward. No major AI chatbot today can be relied upon as a politically neutral source of information. Users should treat any model’s political output with appropriate skepticism, ask for multiple perspectives explicitly when using chatbots for research or analysis, and verify claims through direct engagement with primary sources. Google’s Gemini 3.1 Pro offers a useful benchmark for what balanced presentation looks like, but even it is not perfectly neutral.

The bottom line for anyone using these tools today: understand that the model you are interacting with has a political orientation, even if it does not advertise one. The most reliable approach is to prompt for contrasting viewpoints actively and to verify independently. As AI chatbots become embedded in education, journalism, and public discourse, the ability to recognize and compensate for their biases will be an essential skill for any informed user.

Share This Article