Senior Data Scientist, Safety & Security at Wikimedia Foundation

Location: Remote | Type: Full-time | Category: AI / Machine Learning

Summary Wikipedia is a trusted source of knowledge the world over, read by over a billion people a month in over 300 languages. It’s offered for free and operated independently by the non-profit Wikimedia Foundation (that’s us), powered mainly by small donations. We are hiring a Senior Data Scientist to help guide Wikipedia’s anti-abuse and security strategy as a member of our Product Analytics team. In this role, you will inform our product development strategy by collaborating with product managers, engineers, and others to help our features and interventions make the impact they should.  As the Senior Data Scientist, Safety & Security, you will partner closely with the Product Safety and Integrity team, whose strategy is built around bringing cutting-edge AI and machine learning techniques to bear on difficult problems in the field of safety and security. You will amplify the impact of the safety and security roadmap through creative data analysis, identifying the right metrics for the job, and developing data strategies that advance our product decisions. This role is an opportunity to be at the center of a pivotal time for Wikipedia, as AI changes how people access knowledge and presents new risks and opportunities in online safety and security. Examples of the kinds of projects you may work on include: Setting the strategy for how we measure the success of efforts to detect bots and inauthentic account activity (sockpuppetry, in wiki terms). Providing analysis to inform the development of new heuristics and approaches to detect potentially abusive activity on the platform. Collaborating with research scientists and security engineers to develop and evaluate AI models to find malicious user-authored code. Please note that this role requires availability for critical meetings and synchronous collaboration between 13:00 and 19:00 UTC. You are responsible for: Serving as a strategic partner for product managers, engineers, designers, and other colleagues in product development teams. Proactively making recommendations on product direction and data strategy, surfacing insights and interpreting results. Timely delivery of accurate quantitative data and insights to inform strategy, guide product decisions, and assess the impact of product development work – balancing the need to operate in a fast-changing environment with appropriate rigor. Clearly communicating results and data-backed recommendations to stakeholders across different functions, ensuring understanding and guiding decision-making. Building and executing measurement strategy, evaluating experiments using Bayesian & Frequentist approaches, and performing quantitative research to inform product decisions. Developing queries & code to work with large-scale internal and external data repositories, using cutting-edge AI assistants, and tools such as Apache Iceberg, Hive, Druid, Presto, and Spark. Building automated, self-service dashboards for product teams that track success and health metrics using tools such as DBT, Airflow, and Superset. Helping product team members instrument features, perform data quality assurance, and use tools for running A/B tests. Demonstrating good judgment in prioritizing workload and selecting analytical methods and techniques. Skills and Experience: Experience with experimental design and statistics or machine learning methods. Experience using AI coding assistants for data analysis and other technical tasks. Fluency in R or Python, and common version control and command-line-based tools, for data analysis, statistical modeling, simulation, visualization, & reporting. Fluency in SQL and working with large scale data (we use tools like Hive, Presto, Druid, and Spark). Experience collaborating with product development teams in a fast-paced environment to test, analyze, and evaluate user-facing features for an online platform. Qualities that are important to us: Ability to explain data and insights clearly to non-specialist audiences and gain their understanding and confidence. Curiosity and critical thinking skills; a lifelong learner who sees situations through multiple lenses. Empathy towards and commitment to work with the Wikimedia volunteer communities. Discretion and competence in handling sensitive or confidential data. Commitment to the mission of the organization and our values and guiding principles. Self-motivated with an ability to navigate through ambiguity and complexity. The Wikimedia ecosystem is complex, resources are limited, and our guiding principles are ambitious. We want you to work to find solutions embracing these factors. Additionally, we’d love it if you have: Prior experience working in areas related to trust & safety, information security and privacy, or fraud and anti-abuse. Exposure to, and interest in, ethical data management and privacy practices. Contributed to Wikimedia projects or have experience working in other open source projects. Experience with Super

Related remote job searches

Apply Now

High-intent remote job routes