Trust and Safety (T&S) refers to the organizational functions, teams, policies, and technologies that online platforms use to protect users from harmful content, abusive behavior, fraud, and security threats. The term originated in e-commerce contexts in the late 1990s, where it described efforts to build trust between buyers and sellers in online marketplaces. As social media platforms grew in the 2000s and 2010s, T&S expanded to address challenges related to user-generated content, including harassment, online child safety, hate speech, misinformation, and violent extremism. Trust and Safety work combines human review with automated detection systems to enforce platform policies. The field has faced scrutiny over enforcement practices, labor conditions for moderators, and questions about platform accountability, with regulatory frameworks increasingly mandating specific T&S requirements.
History The concept of "Trust and Safety" (T&S) emerged as a critical function in the early days of e-commerce, driven by the inherent risks of transactions between physically separated and anonymous parties. In a 1999 press release, the online auction site eBay used the term while introducing their "SafeHarbor trust and safety program," which included easy access to escrow services and customer support. Initially, eBay's strategy was built on a "community trust" model as its primary security communication, encouraging users to regulate themselves. "Trust" was used to reference trust among eBay users and between eBay users and eBay itself; "safety" was used to reference keeping platform users safe. With internet platforms growing in scale and complexity, there was a marked increase in the scope of online harms. Social media, app stores, and marketplaces faced unique threats, including impersonation, the spread of malware, and sophisticated scams. The term soon spread throughout the tech industry, expanding from online marketplaces to social media, dating websites, and app stores. Trust and Safety teams emerged as distinct entities across the tech sector, with increasing specialization in fields such as child protection, cyberbullying, and harassment prevention.
Regulatory response While the term evolved within the private sector, judicial rulings began shaping a legislative framework that encouraged investment in online user protection teams. The landmark New York state court decision in Stratton Oakmont, Inc. v. Prodigy Services Co. in 1995 ruled that the early online service provider Prodigy could be liable as a "publisher" for defamatory content posted by a user because the service actively filtered and edited user posts. The ruling created a disincentive for interactive computer services to engage in content moderation, as any attempt to regulate content could result in publisher liability for all user-generated material on their platforms. In response, Congress enacted Section 230 of the Communications Decency Act in 1996, which provided immunity for online platforms hosting third-party content.
Professionalization During the 2010s, the Trust & Safety field matured as a distinct professional discipline with major technology firms investing heavily in Trust and Safety, creating expansive teams staffed by professionals from legal, policy, technical, and social sciences backgrounds. In February 2018, Santa Clara University School of Law hosted the inaugural "Content Moderation & Removal at Scale" conference, where representatives from major technology companies, including Google, Facebook, Reddit, and Pinterest, publicly discussed their content moderation operations for the first time. During this conference, a group of human rights organizations, advocates, and academic experts developed the Santa Clara Principles on Transparency and Accountability in Content Moderation, which outlined standards for meaningful transparency and due process in platform content moderation. The Trust & Safety Professional Association (TSPA) and Trust & Safety Foundation (TSF) were jointly launched in 2020 as a result of the Santa Clara conference. TSPA was formed as a 501(c)(6) non-profit, membership-based organization supporting the global community of professionals working in T&S with resources, peer connections, and spaces for exchanging best practices. TSF is a 501(c)(3) and focuses on research. TSF co-hosted the inaugural Trust & Safety Research Conference held at Stanford University.
Core functions While organizational structures vary across companies, T&S teams typically focus on the following core functions:
Account integrity Account integrity teams work to detect and prevent fraudulent accounts, fake profiles, account takeovers, and coordinated inauthentic behavior. This function combines behavioral analysis, pattern recognition, and authentication systems to identify suspicious account activity, and focuses on ensuring that accounts represent real individuals or legitimate entities and operate within platform guidelines.
Fraud Fraud prevention extends beyond fake accounts to include detection of financial scams, phishing attempts, and marketplace fraud in order to protect users from financial harm and maintains platform trustworthiness for legitimate transactions. This can include manual and automatic detection systems that analyze transaction patterns including velocity (frequency and volume), geographic anomalies, mismatched billing information, and connections to known fraudulent accounts or payment instruments. Machine learning is frequently used to assign risk levels to transactions based on multiple factors, enabling automated blocking of high-risk payments while allowing low-risk transactions to proceed smoothly.
Content moderation Content moderation involves reviewing, classifying, and taking action on content (including user-generated content, advertisements, and company-generated content) that violates a platform's policies, such as explicit nudity, misinformation, graphic violence and any other non compliant materials. Platforms use a combination of automated systems and human reviewers to enforce content policies at scale before the content is live, proactive review after the content is live, and reactive review after the content is live.
… excerpt ends here. Continue reading the full article.
