Algorithmic bias in social media appears when existing algorithms that decide which content people see (posts, videos, or ads) unfairly favor certain groups, opinions, or behaviors over others. Algorithms are designed to learn from data, but the problem with algorithms is that the data they are trained on often reflects human biases or social inequities. If people have historically engaged with some types of content more than others (without the data being representative), then the algorithm can "learn" to take that bias into account when making decisions.
Algorithms are not truly neutral - they reflect human choices that can occur before algorithm training occurs, as well as the data that they are trained on after the fact. Every algorithm is developed by designers, who must pick data, sets of outcomes, and key value metrics indicating success or failure. With these design priorities come assumptions, values, and trade-offs depending, for example, on what "relevant," "engaging," or "credible" means.
The data is not neutral either. It is a record of human behavior. And since human behavior is agonizingly biased (socially, culturally, and even politically) that bias comes right into the algorithm. If people engage with sensational content and click it more or leave comments, then the algorithm will learn to push more sensational content. It will not push it simply because it is "truth," but because it is what will keep you engaged and coming back.
Social media platforms analyze your behavior in order to decide what content to present to you and, furthermore, guess what content will keep you engaged the longest. Every time you "like" a post, share a video, follow someone, or even stop scrolling on the app, you are providing the app with tiny pieces of data on what potentially captures your attention. All that data, in conjunction with information like your location, the time of day, device, and any past interactions, makes up a detailed understanding of your interests and habits. From there, social media platforms rank what content is possible to show you, based on how likely they believe you will engage with the content. Platforms typically desire to maximize interactions to maximize engagement, for example, more "likes", comments, shares, and watch time, because more engagement will keep you in the app longer, and allow the platform to sell ads and generate more revenue. This is why you may recognize your social media feeds getting more personalized over time, the more you remain on the app and engage in its content, the more the platform learns and seeks to improve what it shows you.
Artificial intelligence is an important component of social media algorithms. Many newer social media apps such as Tik Tok, Instagram, YouTube, and Facebook, rely on artificial intelligence to assess your activity on the app and give a customized experience. The artificial intelligence in the background uses machine learning models to analyze vast amounts of user data. It examines your clicks, how long you watched a video, what posts you like, and how often you login. It then uses that data to curate your social feed with the most engaging content based on your interests.
Bias can actually enter the workflow at all three levels: data, model training, and the user engagement feedback loop. But most often bias comes in at the top level with data. The data that is used to train social media algorithms comes from human behaviors, and therefore contains all the already-existing social, cultural, and historical inequalities present. If particular groups are under or misrepresented in that data, the algorithm will learn the same behaviors.
When training a model, bias can be reinforced based on how the AI is programmed to make decisions. Developers must make decisions on the metrics they would like their algorithm to optimize - usually towards engagement like clicks or watch time. That decision creates bias against content that is merely controversial or extreme, because that style of content generates greater responses. The model is learning to chase engagement, not fairness or truth, and that end result is a firmer theme of bias.
And then lastly, there is the user engagement feedback loop which consistently feeds the bias back into the algorithm. When users engage more with specific content, the algorithm interprets that viewing as "success" and will subsequently show similar content to more people. This cycle exponentially amplifies the predisposed bias.
Personalization serves you, while manipulation benefits the system behind the screen. The trick is that personalization and manipulation can merge - what was at first a suggested recommendation can gradually become a kind of control if you never disengage from your algorithmic bubble. That's what makes learning about these systems so important: it will help you recognize when comfort and easy behavior becomes an influence.
Social media companies should indeed be held accountable for what their algorithms do since they can do more than show viral videos, as they affect opinions, feelings, and even democracy. If an AI development company builds an algorithm that spreads false information, promotes hate, or harms the mental wellbeing of users, it cannot only say, “it was the algorithm.” The algorithm is their product, and they profit from it; therefore, they have a duty to ensure it is fair and safe to use.
Data scientists and engineers who design these algorithms also have ethical responsibilities. They are building tools that determine what billions see every day, and therefore need to consider social impact, not just technical performance. This entails asking difficult questions, such as: Is this model perpetuating bias? Is it likely to harm vulnerable communities? Is it sufficiently transparent that the people affected will understand how it impacts them? Ethics in AI is not simply the philosophy of ethics, but part of responsible design.
When it comes to holding platforms accountable, that is no easy task. Most people lack the time, tools or technical understanding to know how the algorithms operate. Regulatory frameworks are often needed to define rules and prevent companies from hiding behind vague assessments of value or claims of uncertainty. Without such external monitoring, the digital environment prioritizes profits ahead of fairness and truth.
Currently, technology companies are incredibly opaque in their operations. Companies offer to some degree qualitative parameters for how their algorithms rank content, but virtually never provide what data is collected or how the information leads to decisions. They typically refer to this information as 'proprietary,' and makes it all but impossible for outsiders to examine if these algorithms process information in a fair way. So, while accountability and transparency are often mentioned, we are far from the day when accountability or transparency will match the power of these algorithms in our society.
There are organizations, tools, and frameworks in explicit ways that promote the reduction of bias. For example, the Algorithmic Justice League is an organization that works to raise awareness and advocate for fair use of AI systems.
Open-source, community-collected fairness toolkits, including AI Fairness 360, provide data scientists with tools to help them diagnose and reduce bias in their datasets and models. After technology, we are now seeing some governance frameworks emerging (such as the Organisation for Economic Co-operation and Development) that support policies around auditing, monitoring, and regulation of algorithmic systems.
Finally, there are more specific domain-initiatives like the United States Equal Employment Opportunity Commission launching work on AI and algorithm use in employment decisions.
Across these three areas, we have technological tools (audit frameworks, fairness toolkits), organizational/policy initiatives (governance, ethics units), and community/advocacy groups all working parallel with evolving bias issues in algorithmic uses, systems, and life..
Independent audits provide a measure of outside review that can uncover biases (in terms of the data , decisions made in model design, and deployment) that may otherwise remain undetected. For instance, the OECD has referenced algorithmic audits being a "cornerstone of AI risk management."
XAI can assist by providing more visibility into the algorithm's logic (such as why a decision was or was not made, what features were most important). This transparency can create the opportunity for stakeholders (data scientists, regulators, users) to examine fairness, or see that something may have gone wrong.
However, audits alone are not a guarantee of fairness. Also, audits may only examine and report on measures of bias, and may miss deeper socio-technical issues (such as power, representation, and context), as articulated through critiques of audit-only approaches.
So they do help to some extent, but audits, like XAI, are simply part of a larger ecosystem of checks, human judgment, diversity, and continuous monitoring.
When platforms grant meaningfully some kind of controls, filters, or scale-back on certain recommendation loops, that puts the user back in the driver's seat of their experience, reducing the feeling of being at the whim of whatever's behind the "black-box" of the algorithm.
Practically speaking, if you can say to the system, "I'm opting out of this kind of content," or "I'd like to see more content in this vein," you can gently nudge your experience out of an echo-chamber or patterns of viewing behavior that can be unpleasant with a recognition of the user instead of the system as the agent, even if we are still stuck with algorithmically generated content.
For this to work, though, the options need to be visible, straightforward to utilize, and of some actual effectiveness (possibly not just cosmetic). There is still the possibility that the algorithm continues to work in the background to deliver unmoderated content that the user isn't aware of, so the user control is symbolic more than real.
Diversity is critically important. When algorithm design teams are comprised of people with diverse identities and backgrounds - such as gender, ethnicity, culture, lived experience - they are more likely to notice assumptions, blind spots, and biases that a homogenous team would overlook. For instance, what may appear to be "normal" data to one group, may embed exclusion to or misrepresent another group.
Creating systems with individuals who hold different perspectives will help the team ask better questions: Are we overlooking certain user populations in our training data set? Are our evaluation metrics related to fairness across populations? Are the features we selected unintentionally disadvantaging someone?
There are small yet concrete steps we can take to diversify our digital systems. For example, we can intentionally pursue a wider variety of voices from varied backgrounds on social media and engage with a wider variety of content with socio-political purposes in mind, thereby dismantling algorithmic bubbles. Viewers can also participate in, promote, and elevate independent and original journalism and other media, utilize and advocate for social media platforms and algorithms geared towards social transparency, and subject content to critical thinking before reading, watching, or sharing with others, which may even mitigate polarized thinking and engagement. Algorithmic fairness is also not absolute in terms of cultural and social values, and we should not think of fairness decisions as being made solely by tech companies, as we should be collectively cultivating fairness in the algorithmic age with input from users, policymakers, ethicists, and so on. Ultimately, fairness in the algorithmic age should be co-created with people, not imposed.