Are large language models widening the digital divide between majority and minority languages? Or can they be harnessed in preserving language diversity?
On Day 3 of Wikimania Paris, representatives of many different language and cultural communities shared their experiences in working within the Wikimedia movement.
Respect for indigenous communities
Erroneous information about indigenous languages and cultures can alienate potential contributors. During the keynote panel, Michelle Collipal, a member of the indigenous Mapuche community in Chile, shared how she has started to motivate Mapudungùn speakers and cultural authorities to engage with the Mapudungùn Wiki project.
Moderator Dr Terri Janke underscored the importance of involving indigenous communities, respecting their knowledge as well as their right to self-determination. What this means in practice for wiki spaces is discussed in the white paper “CutureStrong Platforms: Setting the Standard at Wikimedia”.
Automating and improving translation
During “Language Diversity in the Digital Commons”, Pau Giner of the Wikimedia Foundation shared that the MinT (Machine in Translation) was designed from the start to support diverse language models. The platform currently supports 200 languages.
His ideas for the future include making Wikipedia articles more accessible on mobile phones, and to offer a collection of tools to help people create their first wiki, going beyond translation.
During the Q&A, Kepa Sarasola (User:karasola) of University of Basque reflected on the improvement of translation tools available since he started working on Basque Wikipedia in 2011. The tools, he said, had made it possible to grow minority language Wikipedias – while giving them the flexibility to choose which parts of articles to translate and where to start from scratch. The integration of AI tools has accelerated the process significantly.
Collaboration across communities
The need for collaboration across language communities was a recurring theme across multiple sessions. One example of a cross-border network is Linguatec-IA, an EU project for the digitization of languages in communities in the Pyrenees region, including Basque, Catalan, Occitan, and Aragonese.
David Castillo Parra of UNESCO discussed the launch of the New Commons Incubator for indigenous-led capacity building programs. Applications will be accepted for indigenous-led teams through 14 August.
The need for the Wikimedia movement to remain true to its mission – and to aligning on milestones that matter – was an inspiring message from Audrey Tang in “The State of Wikimedia & AI 2026”. “Don’t let it be a race…Wikipedia never tried to win. We tried to make sure there were many winners at any one time.”
As Jimmy Wales told Le Monde, “It’ll be all right. We’ll adapt, change, use AI in our own way. We’ll find a way.”
For the second time in a row a Ukrainian edition of Wiki Loves Monuments international photo contest had a special category dedicated to Polish heritage in Ukraine — almost 4600 photos by more than 100 authors were submitted. This special category was a joint project of Wikimedia Polska and Wikimedia Ukraine.
A collage of the winning photos of the 2025 Polish Heritage in Ukraine campaign
Mykola Kozlenko (NickK), a member of Wiki Loves Monuments Ukraine organising team, Board member of Wikimedia Ukraine, commented:
“As a background, Wikimedia Ukraine has been organising Wiki Loves Monuments since 2012, with a goal to collect photos of all cultural heritage monuments of Ukraine on Wikimedia Commons, notably for use on Wikimedia projects. One of the main problems we encountered early on was that our official state lists are biased, especially regarding communist heritage, which was the main reason to have monuments protected or not back when Ukraine was under Soviet rule. And it really matters, as participants are more likely to upload pictures of monuments they associate themselves with.
So we started to organise special categories for less represented monuments of national minorities in Ukraine as early as 2013 — Armenian, Greek, later Crimean Tatar, Jewish, German, Polish, and the most recent one, Bulgarian. It does require more work from us, as we also need to find information about monuments that are not listed officially, or read additional sources to add this or that object to the special categories lists. But it helps us to expand the database of cultural heritage, make it less biased. And it helps us to increase awareness, and motivate veteran participants to continue taking part in the contest, and attract new participants, either interested in the multicultural past of Ukraine, or being a part of those national minorities themselves.
The Polish Heritage in Ukraine campaign was conducted for the second time, and we are very pleased with the results — almost 4600 pictures uploaded by more than 100 authors, depicting 373 monuments, and out of them — 41 are not officially protected by the state, so they are even more endangered, as they can be not only destroyed or damaged by russian drones or rockets, but they can be demolished or repurposed with no oversight by the cultural heritage protection authorities.
The main purpose of a separate special category continues to be to draw attention to these monuments and their condition, and to document them for Wikipedia. And we are very grateful for the support of Wikimedia Polska, that made this project possible, and also to our volunteers and participants for their continued active involvement in the project”.
The Polish Heritage in Ukraine campaign was happening alongside the main contest period for Ukraine in October 2025. Volunteers updated the lists for the special category, so as of now it is containing 1358 monuments (488 out of them with no official protective status). During the campaign itself 102 participants submitted almost 4600 photos, picturing 373 monuments (41 out of them are not registered as monuments) from 15 regions of Ukraine. 24 monuments were pictured for the first time.
A Wiki Loves Monuments Ukraine barnstar
Due to a considerable number of submitted works, there was a pre-selection round. 16 volunteers from Poland took part in reviewing the images. Some Polish volunteers shared their reflections on the photos, their motivation to help, and the process.
Piotr “PMG” Gackowski, Polish volunteer helping with preselection, an editor with almost 9 mln edits on Wikimedia Commons, a Polish Wikipedia administrator, reflected:
“I participated in the preselection of photos for many reasons. One of them is patriotism. In this way, I can support the memory of Poland and the Polish people. The second point is the curiosity typical of every Wikipedian: I took part in the Polish WikiLovesMonuments and wanted to see what photographs from other countries look like. For me, the difference was that the photographs from Ukraine that I was rating much more frequently showed objects in rural areas. In Poland, large cities dominate, so in my opinion it was a significant difference. At the same time, it is important to me that I can help Wikipedians from Ukraine in their work. I am aware that every monument they commemorate by taking photographs could be destroyed”.
Teukros, a Polish pre-selection volunteer, and a Polish Wikipedia administrator, commented:
“I have participated in the photo preselection process for the Wiki Loves Monuments campaign (Polskie Dziedzictwo w Ukrainie) twice now, and I have genuinely enjoyed doing so. To be honest, I did not have any particularly special reasons for joining this initiative – the simple fact that the Wikimedia community in Ukraine had asked for assistance was reason enough for me.
My experience of participating has been a mixture of sadness and joy. Sadness, because it is plainly visible that Polish heritage sites in Ukraine are often damaged, neglected, and that there is little indication that this situation will improve in the near future. Joy, because the very fact that the Ukrainian community has taken on such a challenge allows us to believe that at least the memory of the Polish presence in these lands will endure.
If I were to say what inspired the greatest sympathy in me during this project, it would, paradoxically, be the photographs that I had to reject. Crooked, overexposed, blurry, taken by amateurs without any special preparation – they were perhaps the strongest testimony that, among completely ordinary people, the memory of the shared history of Poland and Ukraine is still very much alive”.
The pre-selection volunteers reviewed 4571 images (the organising team removed images submitted by the participants with conflict of interest, like organisers and jury members), divided in such a way, that each image was viewed by 3 volunteers.
Archiwald, a Polish pre-selection volunteer, and a Polish Wikipedia administrator, commented:
“In February of this year, I received an offer through WMPL to participate in the preliminary selection process as a person assisting with the initial evaluation of photos. I was happy to join the effort, especially since I already had some experience with similar initiatives at the time. As an editor who focuses, among other things, on historical matters, I realize just how useful the files I’ve been reviewing will be. I’m not just referring to the winning photos here; even those that ultimately didn’t receive any awards add significant value to the Commons resources.
It’s a very pleasant feeling to look through the contest results and notice instances where the judges rated a photo just as highly as I had earlier. I felt that way, for example, when I noticed that the photographs of the palace in Pryozerne by Oleksandr Malyon had been recognized. Photographs like these have immense historical value, which usually becomes apparent only after many years; therefore, the author’s decision to make them available under free licenses deserves recognition”.
Adrian Tync (Gower), another Polish pre-selection volunteer, active on Wikidata, Wikimedia Commons, and Polish Wikipedia, shared:
“I got involved in the photo pre-selection process because I’d taken part in the ‘Wiki Loves Monuments’ competition a few times myself as a photographer, and I was curious to see what it was like from the other side. I enjoy browsing and evaluating other people’s photos on Commons, for example in the Quality Images nominees section, so this was the perfect task for me. I’m interested in Polish historical monuments and Polish cultural heritage, and thanks to the pre-selection process, I got to see many of them”.
Out of this round 815 photos proceeded to the next round. The organisers reviewed the images more closely, and removed the ones that were not up to the technical standards (like lower resolution), so 711 images proceeded to the main jury, which included Polish Wikimedians and partners of Wikimedia Polska:
Damian Kujawa — Wikimedian, volunteer, activist;
Julia Szablowska — photo editor, photographer, curator;
Magdalena Lachowicz — Assistant Professor at the Department of Eastern Studies at Adam Mickiewicz University in Poznań, Poland;
Each work was viewed by all three jury members. 120 pictures made it to round two, where each jury member was asked to evaluate each work from 1 (minimum) to 10 (maximum) points. The guidance when evaluating pictures was:
from 0 up to 3 for technical quality (sharpness, use of light, perspective etc.);
from 0 to 3 for usefulness of the image for Wikipedia;
from 0 to 3 for originality.
1 additional point for something special in the picture.
The results are presented below, and they are grouped thematically, to showcase the breadth and depth of Polish Heritage in Ukraine, so the awarded works are from different regions of Ukraine, and are grouped by different objects depicted (churches, castles etc). No separate award for active participation — people awarded are among active contributors.
Best photos – Churches (pol. Najlepsze fotografie – Kościoły)
Holy Trinity Church (2022). Velykyi Ostrozhok, Vinnytsia Oblast
The author uploaded the first ever pictures not only of the church, but even from the village itself. And his pictures are now illustrating the article about the village on Wikidata and local Wikipedias (Wielki Ostróżek in Polish Wikipedia, for example). The version in Ukrainian did not even contain the mention of the church, as the building is not a listed monument officially.
Best photos – Palaces, Estates (pol. Najlepsze fotografie – Pałace, Majątki)
Potocki Palace (2025). Tulchun, Vinnytsia Oblast
Rej manor (2025). Pryozerne, Ivano-Frankivsk
Best photos – Other Buildings (pol. Najlepsze fotografie – Inne budynki)
The building was built by a Polish architect Mikołaj Tołwiński, there is no article about the school itself on Polish Wikipedia yet. Due to the Russian occupation of Mariupol, getting new free pictures (or even up to date information about the state of the building) is not going to be a trivial task.
Former house of the Branicki estate manager (2022). Rozkishna, Kyiv Oblast
This is also not a listed building, which makes its status to be more endangered — it is now privately owned, and there were news about it being on sale.
Best photos – Chapels (pol. Najlepsze fotografie – Kaplice)
The results of the joint project and winners were celebrated at the Wiki Loves Monuments Ukraine hybrid awards ceremony on May 30, 2026.
At the 2025 Wiki Loves Monuments Ukraine Awards Ceremony
Polish heritage in Ukraine Special category statistics
Winners present offline
Olena Suhak, one of the winners, commenting on her works
Serhii Plakhotniuk, one of the winners, commenting on his photo
Volodymyr Tarasov, one of the winners, commenting on his contributions
Oleksandr Malyon (on screen), one of the winners
The winners absent at the event will receive their prizes by post.
Iryna Boiko, communications manager of Wikimedia Ukraine, commented:
“This is the second time we are organising the “Polish Heritage” special category in the Ukrainian edition of the Wiki Lobes Monuments contest. The request to organise a separate category for Polish sites has been repeatedly expressed by the participants themselves, because many such sites are now falling into disrepair and simply collapsing in the absence of an active community to care for them.
The main focus of the special category was churches, Polish cemeteries, castles and fortresses, but the jury also paid attention to residential and other buildings. Of course, the parameters of inclusion in the contest lists are quite wide, because the very idea of the “Wiki Loves Monuments” competition is to collect photos to illustrate Wikipedia articles, and in order for the article to illustrate the life of a community or a certain period, the parameters of inclusion should be quite wide. So, the special category lists contain buildings created by Polish architects for Polish activists, or buildings where Poles lived, or buildings that were important for the Polish community of a particular settlement — like a bookstore in Kharkiv, that became a center of a local Polish community life.
My personal favorite photo, a very symbolic embodiment of what we are trying to achieve through this joint project with Wikimedia Polska, was the work of Valentyn Mahovkin, our long-time participant, which depicts the process of restoring an epitaph on a tombstone in a Polish cemetery in the village of Chornyi Ostriv in Khmelnytskyi Oblast. And, by the way, this cemetery is not officially protected, and only its gate has an official status as a cultural monument…”
Restoration of the epitaph (2024). Polish cemetery. Chornyi Ostriv, Khmelnytskyi Oblast
Wikiquote training in Bandung (Hasnanf, CC BY-SA 4.0 via Wikimedia Commons)
When people think about Wikimedia projects, most only know the world’s largest online encyclopedia, Wikipedia. Many do not know that Wikipedia has dozens of sister projects, including Wikimedia Commons, Wiktionary, Wikibooks, Wikisource, and Wikiquote. One of the lesser-known projects is Wikiquote. It is a collaborative project that collects and preserves notable quotations from famous people, films or series, fiction and non-fiction books, proverbs, and well-known sayings.
In Indonesia, Wikiquote is currently available in three languages, Sundanese, Banjar, and Gorontalo. In Wikimedia Bandung Community, we believe that community collaboration is essential to enriching Wikiquote’s content and introducing the project to more people. With this goal in mind, we launched WikiSuarana, a community initiative to enrich the Sundanese Wikiquote with quotations related to memorable events and popular trends from 2025.
Project outcomes
WikiSuarana project was carried out by four members of Wikimedia Bandung community: Hasnanf, Raflinoer32, Zulaihamaryam, and Sonofbrahma from March to May 2026. Together, we created 224 Wikiquote articles featuring quotations related to events that took place throughout 2025. Each team member contributed 56 articles to Sundanese Wikiquote.
In addition to our work on Wikiquote, we also contributed to Sundanese Wikipedia by creating 56 articles about notable events from 2025. We know that many Sundanese speakers still use Sundanese Wikipedia as a source of information. By enriching Wikipedia with these articles and linking them to the related Wikiquote pages, we make it easier for readers to discover and explore our collection of quotations on Sundanese Wikiquote.
Community outreach through training and meets up
Community outreach for Wikisuarana was carried out during the month of April 2026, consisting of a series of training and meet-up activities. As a warm-up activity, in the first week of April, we conducted a community meet-up to edit on Sundanese Wikipedia together focusing on creating new articles regarding remarkable events that happened throughout the year 2025 and the notable figures related to it. This event was attended by 13 participants, both online and offline, which created 14 new articles on Sundanese Wikipedia. By starting WikiSuarana with this thematic edit activity, it was expected to provide a thematic context regarding the purpose of this project, which is to document remarkable events and notable figures in Sundanese Wiki projects. Some of the articles that were the results of this event, for example Tambang Grasberg, Satelit Nusantara Lima, and Sri Rejeki Isman.
On April 18, the Asia-Africa Conference is annually commemorated in Bandung, since the city was the first host of the conference back then in 1955. On that day this year, we conducted a Wikiquote training activity partnering with an independent library in Bandung, which carried out the theme of anti-colonialism spirit. In this activity, we focused on training new contributors to edit and create new quotation articles on Sundanese Wikiquote regarding anti-colonial figures, anti-colonial literatures, and those related to the Asia-Africa Conference. The training was attended by 13 participants, which created 15 new articles. Hopefully, other than attracting new contributors, this event could be the beginning of a consistent effort to amplify and document the voice of anti-colonial figures through Sundanese Wikiquote. Some of the articles that were the results of this event, for example Behind The Scenes – Story of The Bandung Conference Committee, Teh dan Pengkhianat, and Leila Khaled.
The series of WikiSuarana was concluded with another community meet-up as a continuation of the prior training, which was focused to create and edit quotation articles on Sundanese Wikiquote regarding the theme of anti-colonialism and remarkable events that happened during the year 2025. This event was attended by 11 participants, both online and offline, which was also attended by some of the participants of the previous training event, and produced 14 new articles. Some of the articles that were the results of this event, for example Jawaharlal Nehru, Mohammad Yamin, and Ali Sastroamidjojo.
Lesson learned
Through WikiSuarana project, we all learned two valuable lessons:
The first is the importance of partnerships. As a local community, collaborating with organizations that share our mission is essential to promoting free knowledge, especially about the Sundanese language and culture. These partnerships help introduce Wikimedia projects and our community to a wider audience. During this project, we collaborated with an independent library in Bandung. In the future, we hope to work with more partners to organize Wikimedia activities such as workshops, research projects, and community meetups.
The second lesson is about promotion and outreach. In today’s digital world, many people get information through social media platforms such as Instagram and TikTok. Throughout the project, we created promotional content, from the project launch to activity announcements. However, we learned that relying on just one or two social media platforms is not enough. Recently, Threads has become increasingly popular, and one of our team members found that project posters shared there reached a wider audience and received positive engagement. Based on this experience, Wikimedia Bandung plans to use Threads alongside our other social media channels to promote future Wikimedia activities.
What’s next?
Documenting and amplifying the voices of the people who are part of history is a way to preserve our collective memory. By narrating them in their own voices, hopefully we can make sure that the history being told is honest-to-goodness. Through contributing it to Wiki projects, especially Sundanese Wikiquote, we also hope to be able to preserve them in our mother tongue.
Our plan forward is to keep documenting other voices that are still unheard while also promoting the sister projects of Wikipedia, which already has the Sundanese version, such as Wikiquote. We also plan to reach outward to other parts of West Java to attract many other contributors so that our community can keep growing while also adding much other knowledge to the Wiki projects itself.
I started working as a volunteer community manager for Wiki for Human Rights, Nigeria, in November 2025 with a focus on organizing trainings and retaining LGBTIQ+ Nigerian editors as active contributors to Wikimedia projects. Although I created my Wikimedia account on 1 March 2024, I did not understand how editing worked and never made any contributions after creating it.
The same year, I attended a Wikimedia session on Queerpedia. As a writer who is passionate about volunteering, particularly in the open knowledge movement, I still left without knowing how to contribute. The session ended with participants creating accounts but without a practical understanding of how Wikimedia actually worked. This is something I have observed among newcomers: navigating Wikipedia, the most well-known Wikimedia project, can be a daunting experience.
After-session group picture
That changed when I attended Wiki Loves Pride 2025, held on 29 June. It was my first in-person Wikimedia event, and with my laptop beside me, everything about contributing to Wikipedia suddenly became much clearer. Through editing Wikipedia, I also discovered the wider Wikimedia ecosystem and its sister projects. Even after the training, I still encountered a few challenges, particularly with adding awards, information tables, and infoboxes. However, through conversations on WhatsApp with the Wikimedia Nigeria Project Officer, Ayokanmi Oyeyemi (user: Kaizenify), who facilitated the training, as well as guidance from Wikipedia help pages and Wikimedia Commons documentation, I was able to overcome those challenges.
Since assuming the role of Community Manager and Project Officer for Wiki for Human Rights, Nigeria, I organised monthly virtual and physical training sessions aimed at improving editor retention. Organising both the physical and virtual Wiki Loves Pride campaign this June felt like a full-circle moment; déjà vu. It also became an opportunity to reflect on the learning experiences from reviewing participants’ contributions and identifying areas where new editors commonly struggled.
Tony Obinna facilitating a session
Since becoming an active Wikimedia contributor in July 2025, I have made more than 3,000 edits across Wikimedia projects, with a primary focus on LGBTIQ+ and Nigerian topics. Beyond editing, I have also taken on leadership positions, including serving as a committee member for the Nigerian National Funding program and as a core organizing member for Queering Wiki, scheduled to take place later this year in Canada.
Alongside the online Wiki Loves Pride campaign, we partnered with the Centre for Population Health Initiatives (CPHI), a health organization that serves both the general population and minority communities, to host Wiki Loves Pride. The program combined a Pride celebration with Wikimedia training for both experienced and new editors. It also marked the first time I independently facilitated an entire Wikimedia training session from the beginning to the end: account creation, making edits, and introducing participants to the broader Wikimedia ecosystem.
Participants engaged throughout the session
As someone who has always dreaded public speaking because of a minor speech impediment, becoming part of the Wikimedia movement as a community leader has helped me grow tremendously. Standing in front of more than 15 participants and leading a session on documenting queer knowledge, I did not freeze or lose confidence. Community advocacy for LGBTIQ+ people in Nigeria has always required me to speak publicly from time to time, but Wikimedia has made it a consistent part of my work through monthly virtual and physical trainings. It has taught me that confidence in public speaking is often built through practice, and that many of the fears we carry can gradually be overcome through repeated experience.
When the Wiki Afrodemics Mentorship Programme kicked off in March 2026, I knew I was stepping into something special, what started as a curiosity to learn more about Wikimedia projects has turned into a transformative journey that has completely reshaped how I contribute to free knowledge.
The Wiki Afrodemics Mentorship Programme assembled 20 passionate participants from underrepresented African countries, super proud to be one of the selected participants from the pool of over 300 applicants. This project empower fellows through structured training sessions, hands-on editing, and collaborative activities across Wikipedia, Wikidata, and other Wikipedia sister projects. Being part of this diverse community of learners was inspiring, we came from different countries, spoke different languages, but shared a common mission.
Wiki afrodemics and mentorships programme mentees
Enhancing My Wikidata Skills
A personal highlight of this programme was the significant improvement in my Wikidata editing skills, made possible through the exceptional facilitation of my favourite mentor, David Partey. While I was already familiar with Wikidata , Ialways believe there’s more to improve on; structured data requires a distinct mindset and technical approach.
David’s training sessions were invaluable. He broke down complex concepts like creating new items, adding statements with reliable references, and querying data from the Wikidata platform. His patient and structured approach demystified Wikidata, turning it from a daunting database into an intuitive and powerful tool for enhancing the visibility of African academics.
Under his guidance, I learned not just how to edit Wikidata, but why it matters. I now understand how to;
Add meaningful statements with reliable references
Connect Wikidata to Wikipedia articles and Wikimedia Commons files
Use Wikidata to make African academics more visible online.
This newfound proficiency has made me a more confident and well-rounded editor, capable of contributing meaningfully across multiple Wikimedia projects. Today, I can confidently say that my Wikidata editing skills have improved tremendously, and I owe so much of that growth to David’s exceptional facilitation.
The sessions facilitated by other mentors were equally impactful. Each mentor brought unique expertise and perspectives, and I soaked up every bit of knowledge they shared. The collaborative atmosphere, the peer feedback, and the sense of community made learning feel less like a classroom and more like a family gathering.
Up next!
As this maiden Cohort wraps up, I am filled with so much gratitude for the mentors who invested their time and expertise in us, for the facilitator who believed in this vision, and for my fellow fellows who made this journey so memorable.
The Wiki Afrodemics Mentorship Programme has shown me that mentorship is not just about receiving, it is about growing, connecting, and ultimately giving back. I am leaving this programme as a Wikipedian, a more knowledgeable Wikidata contributor, and a passionate advocate for free knowledge in Africa.
I cannot wait to apply everything I have learned and to contribute to future cohorts, this time, not as a fellow, but as someone who can support and inspire others just as I was supported and inspired.
Thank you, Wiki Afrodemics, for this life-changing opportunity. This is just the beginning of my journey.
How can the Wikimedia community defend freedom, equity, and reliability on the internet? Are we in a global information crisis?
Several sessions on Day 2 of Wikimania Paris explored these questions from different angles. The morning kicked off with breakout sessions on how global trends are impacting government regulation – with attendees joining Wikimedia Foundation board members in small group discussions. Protecting free knowledge was a core theme throughout the day.
Many readers have also shifted from reading about climate change to reading more about other issues such as the cost of living crisis. Dr. Femke Nijsse (User:Femke) discussed the importance of meeting readers where they are – and explaining the science related to these issues.
While AI-generated content may have the sheen of reliability, it often turns out that the sources they cite are hallucinations or that they do not actually verify the claims made by the LLM. Editors on Wikipedia are now starting to use tools such as AI Source Verification (which itself uses LLMs) to predict the verifiability of claims made in a article.
AI crawlers: Encroaching on creativity?
“Collateral Damage? Human Creativity and Interaction in the AI Crawling Era” explored how organizations in the free knowledge ecosystem are responding to the massive increase in AI scrapers.
There was a consensus among panelists that attribution is a critical concern for authors. Creative Commons CEO Anna Turnadóttir and Monica Westin of Cambridge University Press discussed the need to educate authors about the benefits of open access models – while also acknowledging the need to experiment.
Mark Graham discussed how the Internet Archive is reaching out to news organizations to discuss alternatives to blocking the Wayback Machine, such as rate limiting and allowing access only for certain uses.
Striking a balance
Throughout the day, speakers debated the complexities of regulation – and how the rush to “do something” can backfire. Turndóttir said, “What used to be the internet handshake online is now the middle finger. And that sort of environment forces lawmakers to reach for blunt tools like regulation. The community needs to establish norms, but sometimes regulation does cause real harm.”
The keynote session on “Protecting Free Knowledge – The New Battlegrounds of Digital Freedom”, Nnenna Nwakanma (from the internet) discussed the discourse around regulation – and how European approaches may not be applicable in Africa. Panelists during the session, including Axelle Lemaire, architect of the 2016 loi numerique, emphasized the need for balance in protecting openness and freedom online, while also protecting the privacy of individuals.
The conference is making waves in France with media coverage in 25 outlets – including La Croix‘s print edition, a Radio France podcast, and more.
Group photo of the Indic Wikimedia Hackathon Hyderabad 2026, Image by Nivas
Program Purpose
Fifty-six contributors gathered at IIIT Hyderabad for three days to improve Wikimedia’s technical ecosystem. Unlike traditional hackathons that focus primarily on rapid prototyping, the Indic Wikimedia Hackathon 2026 introduced dedicated refinement sessions that encouraged participants to improve code quality, documentation, usability, and long-term maintainability.
Wikimedia hackathons are spaces for developers, designers, content editors, and other community stakeholders to collaborate on building technical solutions that help improve tools, workflows, and overall user experience across Wikimedia projects.
This hackathon is designed for:
Technical contributors active in the Wikimedia technical ecosystem, which includes developers, maintainers (admins/interface admins), translators, designers, researchers, documentation writers, etc.
Content contributors having an in-depth understanding of technical issues in their Wikimedia projects, like Wikipedia, Wikisource, Wiktionary, etc.
Contributors to any other open-source community or those who have participated in Wikimedia events in the past, and would like to get started with contributing to Wikimedia technical spaces.
Participants worked on a curated set of technical tasks prepared by mentors and organizers. They were also encouraged to propose their own project ideas, provided they included a clear problem statement, implementation approach, and were reviewed by mentors before the event.
Building on the experience and learnings from previous hackathons, this event was more efficient, inclusive, and collaborative.
The event aimed to involve more developers who have experience with the Wikimedia ecosystem and had prior experience already contributing to tools, extensions, gadgets, or other technical projects. Editors were paired up with developers to provide domain knowledge, helping them better understand editing workflows, user needs, and the intended behaviour of the applications/ extensions/ gadgets being developed.
Unlike other hackathons where rapid development is the primary focus, this event was not completely hacking but also incorporated dedicated refinement sessions.These sessions encouraged participants to improve the quality of the works by refining design, UI, data privacy, code optimization, documentation and overall maintainability. Additionally, the program also included brainstorming sessions, group discussions, workshops to help participants understand a broader perspective of this ecosystem beyond their individual projects.
Scope and Timeline
The hackathon was conducted as a three day in-person event at the International Institute of Information Technology, Hyderabad (IIIT-H), a long-standing partner that provides space for technical and community events. The venue supported collaborative work through dedicated hacking spaces, mentor interactions, and discussion areas.
The scope of the event was flexible enough to encourage participants to work on a curated set of Wikimedia-related technical projects and tasks suitable for a hackathon which prepared by mentors or a custom project proposed by their own and reviewed by experienced developers and organizers, along with Team Challenges from Wikimania Hackathon 2026.
The program was structured into distinct phases. The first half of the event (approximately one and a half days) focused entirely on development where the first part of the first day was catered to some introductions and welcome notes, ground rules, ice breaker activities, followed by continuous hacking. On the second day, the event had a social activity and the morning session was mostly catered to workshops, brainstorming sessions, and getting to know about OKI work. The second half started with refinement phases. The third day focused completely on refinement, wrap-up and showcase.
Attendance
Total Attendees: 56
Organisers: 11
Mentors: 15
Participants: 25
Editors: 5
The majority of participants were developers with prior experience in the Wikimedia technical ecosystem. A smaller group consisted of experienced Wikimedia editors with technical knowledge, who collaborated with developers by providing domain expertise and user perspectives during the hackathon.
Activities Conducted
Prior to the hackathon, an orientation call was conducted for participants to give an overview of the program, the Wikimedia technical ecosystem, team formation details, and some logistical and operational arrangements. The session also introduced participants to Wikimedia, its technical ecosystem, guided them through basic account setup, and explained the overall hackathon format.
Following the orientation, participants were encouraged to have a call with their specific teams and mentors. These discussions helped participants better understand their assigned projects, including the scope, objectives, expected outcomes, and technical requirements. Mentors introduced project-specific workflows, outlining the scope, objectives, and tasks for each participant to help them engage effectively during the hackathon.
During the event, participants worked on pre-curated tasks across multiple Wikimedia-related projects and collaborated closely with mentors to understand issue tracking, patch submission, and debugging workflows. Mentors supported participants across different projects, helping them navigate both technical challenges and Wikimedia-specific contribution processes.
To complement the technical program, the hackathon also included community-building activities. An icebreaker session at the beginning of the event helped participants interact and build connections. On the second day, interested participants joined a social walk around Hyderabad, providing an informal opportunity for networking. A dedicated women’s dinner was also organized to foster stronger connections among women participants, encourage inclusion, and support long-term retention within the Wikimedia technical community.
Towards the conclusion of the event, a project showcase and presentations session was conducted, during which participants demonstrated their work and shared learnings with fellow participants, mentors, and organizers.
Outputs and Outcomes
The hackathon enabled participants to work on pre-identified tasks across multiple Wikimedia-related repositories, resulting in code contributions, feature enhancements, bug fixes, and documentation improvements. Throughout the event, participants gained practical experience with Wikimedia development workflows, including issue tracking, patch submission, code review, and collaborative problem-solving, with guidance from experienced mentors. Several participants continued engaging with their assigned projects after the event, indicating effective onboarding into Wikimedia technical workflows.
During the hackathon, participants submitted a total of 56 Phabricator tickets across multiple Wikimedia-related projects. The distribution of contributions is summarised below:
Clip2Commons: 9
Deployr: 7
Language Selector Rewrite: 1
Lingua Libre: 5
Montage: 4
NPOV Drift Detector: 1
Observability Tool: 7
Onboarding of New Wikipedia Editors : 2
OpenSpeaks Subtitler: 5
OpenSpeaks Tome: 3
Pywikibot : 3
Scribe: 6
Translate Tagger: 8
ULS Extension :8
Wanda Extension: 4
WikiEval Tool : 3
Wikievol Tool : 3
WikiLinkua : 2
Wikimedia Commons Android: 4
Wikisource Reader App: 2
Total: 87 Repo/Phabricator tickets linked
As part of the Indic Wikimedia Hackathon Hyderabad 2026, several teams worked on projects that directly align with the official Team Challenges announced for Wikimania 2026.
Lingua Libre – Boost multimedia experience, Connect multilingual knowledge
WikiLinkua – Gamify knowledge, Stream data with Wikidata
Wiki Translate Tagger – Connect multilingual knowledge
Wanda / WandaScore / WandaScribe – Welcoming newcomers, Fix the sources / Update the obsolete, The editor of the future
Scribe – Stream data with Wikidata, Connect multilingual knowledge
WikiEvolution – Explore knowledge
WikiNPOV Drift Detector – Deciphering biases
Language selector rewrite – Connect multilingual knowledge
These contributions included code changes, improvements, and related updates submitted under mentor guidance.
What went well:
Program design:
The overall event design enabled participants to collaborate with peers from diverse backgrounds and work effectively on technical projects.
The combination of structured onboarding, continuous hacking, and dedicated refinement sessions supported steady progress throughout the event.
Refinement sessions encouraged participants to improve code quality, documentation, user interface, and maintainability rather than focusing solely on completing tasks.
Mentorship and technical contributions:
Clear project introductions and continuous mentor support helped participants engage confidently with their assigned tasks.
Most of the identified hackathon tasks were actively worked on during the event.
Several participants continued contributing to their assigned projects after the hackathon, demonstrating successful onboarding into Wikimedia technical workflows.
Collaboration:
Pairing editors with developers proved valuable, as editors helped developers better understand user workflows, expected tool behaviour, and usability considerations.
Workshops and discussion sessions complemented the hacking sessions by providing participants with a broader understanding of the Wikimedia technical ecosystem.
Diversity and inclusion:
The hackathon achieved approximately <>% women participation, the highest among events organized by the User Group to date.
The women’s dinner helped foster stronger connections among women participants and contributed to a more welcoming environment.
What can be improved/Learnings?
Program schedule
Participants requested additional time for the project showcase and presentations.
Fifteen-minute breaks were considered too short and could be extended in future editions.
Starting sessions at 9:00 AM posed challenges for some participants because of commuting time.
Project selection
Hackathon tasks require more review before the event to reduce duplication of effort.
Greater emphasis should be placed on improving existing Wikimedia tools rather than developing new ones where similar solutions already exist.
Mentorship
Participants experienced delays when waiting for scheduled online mentor support.
Increasing mentor availability or ensuring more mentors are physically present during the event could improve the overall experience.
What’s next:
Based on the outcomes and observations from the Indic Wikimedia Hackathon Hyderabad 2026, the following recommendations are proposed for future editions of the event:
Prioritise improving existing Wikimedia tools and applications over developing new ones, where appropriate.
Continue incorporating dedicated refinement sessions to improve the quality and sustainability of contributions.
Allocate more time for project showcases, presentations, and participant discussions
Review the event schedule by extending break durations and considering a later start time where feasible.
Continue initiatives that promote diversity and inclusion, including activities that support participation and retention of women contributors.
Strengthen post-event follow-up and mentorship to encourage continued contributions beyond the hackathon.
Explore organizing additional hackathons, workshops, and technical events in college campuses to improve accessibility and encourage new technical contributors.
When people think about hackathons, they often imagine coding sessions, project demos, and late nights spent debugging. For me, Wikimedia Hackathon 2026 in Milan was something much more meaningful: a reminder that some of the most impactful ideas begin as simple conversations between passionate people.
Like many Wikimedia Commons contributors, I have often found it difficult to discover media through traditional search. Commons hosts millions of images, videos, and audio files, but today’s search relies mostly on filenames, categories, descriptions, and structured data not on what is actually visible or audible in the file itself. Hundreds of campaigns, contests and content initiatives add new media every year, which only makes the problem bigger. I kept wondering, half as a joke and half seriously: what if we could search Commons based on what an image or video actually shows, rather than what someone happened to type into its metadata?
It felt like a wild idea at the time, so I shared it with the Wikimedia technical community mostly to see if anyone else found it interesting.
Eugene, David and Gopa at WMHACK 2026
A researcher and developer named David, who had recently joined the Wikimedia community, came across the discussion. David had been working on a technology called WISE.. at the University of Oxford research focused on semantic understanding of visual and multimedia content, enabling people to search images, videos, and audio using natural language. What started as a random online exchange quickly turned into a collaboration. David hadn’t originally planned to attend the hackathon in Milan, but as we kept talking, we both got excited about bringing this kind of search to the Wikimedia ecosystem, and decided to meet in person and try to build it.
Looking back, that decision changed everything.
For two intense days at the hackathon, we worked side by side to integrate and demonstrate WISE for Wikimedia Commons — architecture, datasets, search quality, the usual string of small technical fires. One moment that stuck with me: the first time we typed “horse in an airplane” into the prototype, half-expecting nothing, and watched it actually return the right video. That was the point it stopped feeling like a demo and started feeling like a real tool.
What inspired me as much as the technology was David’s approach to problem-solving. In an era where most developers reach for an AI assistant the moment something breaks, David would just as often dive into technical documentation, research papers, and manuals first. Watching him methodically work through a problem rather than skip to an answer was a good reminder that strong engineering fundamentals haven’t gone anywhere.
What we built
The result of those two days is WISE, a new experimental search experience for Wikimedia Commons. It currently indexes Media of the Day content around 5,000 videos and is built to understand the actual visual and audio content of a file rather than its metadata.
Project Wise showcasing Semantic Image Search
Semantic search. Search using natural language and find media based on what appears in the image or video itself. A few examples that worked surprisingly well:
“man at a train station”
“horse in an airplane”
“man with a flower”
“pirate with a pistol”
Audio search. Search within audio content to find relevant recordings and segments.
Face search. Upload a photo of a face, and WISE can locate that person across images and videos… for video, it can even surface the timestamps where they appear.
Multilingual search. Queries work in multiple languages, including Hindi and Telugu. This matters more than it might first appear: most of Commons’ existing search tooling is built around English-language metadata, which quietly shuts out a large share of Wikimedia’s global, non-English-speaking contributor and reader base. A search experience that understands a query in Hindi or Telugu as well as it understands one in English is a small but real step toward making Commons more usable for the movement it actually serves.
This is only the beginning. We’re already discussing:
Expanding indexing beyond Media of the Day to cover Commons at a much larger scale.
Finding visually similar images after an upload.
Suggesting categories, filenames, and metadata based on visual similarity.
Improving search quality and broadening multilingual support further.
Most of all, this experience reminded me why I love being part of the Wikimedia movement. A passing idea shared on a community forum connected two people from different backgrounds and different parts of the world. An online discussion became an in-person collaboration. A concept became a working prototype. And a hackathon became the place that vision came to life.
Sometimes the most valuable outcome of sharing an idea isn’t the idea itself it’s the people who find it, connect with it and decide to build something together. For me, WISE for Commons is more than a search tool. It’s proof of that.
When I applied for Train the Trainer (TTT) 2026, I expected to learn how to organize better events, facilitate workshops, and become a more effective trainer.
Over three days in Hyderabad, I certainly learned those skills but I also came away with something far more valuable: a new understanding of how strong Wikimedia communities are built and sustained.
Rather than focusing only on editing or technical skills, the program explored the people behind Wikimedia the contributors, organizers, mentors, and volunteers who make free knowledge possible. Through discussions, hands-on activities, and collaborative exercises, TTT encouraged participants to think beyond individual contributions and toward building welcoming, resilient communities.
Here are some of the lessons that stayed with me long after the program ended.
Communities come before content
One of the biggest surprises on the first day was that very little time was spent talking about editing, and all.
Instead, sessions explored trust, belonging, leadership, contributor motivation, and community participation.
The panel discussion “What Makes Us Stay? Trust and Participation in Communities” highlighted something every Wikimedia community experiences: attracting contributors is only the first step. Helping people feel welcomed, supported, and valued is what encourages them to stay.
Another activity challenged participants to analyze contributor data across language communities. Looking at the gap between the large number of people who consume knowledge online and the relatively small number who actively contribute made me rethink community growth. I realized that successful outreach is not measured only by how many people join an event it is also measured by how many continue contributing afterwards.
The sessions on leadership, trust, and the Universal Code of Conduct reinforced another important message: healthy communities are built intentionally. Trust is earned through respectful collaboration, inclusive spaces, and consistent support for newcomers.
Wikimedia Commons is about preserving knowledge not just photographs
Before attending TTT, I believed I already understood Wikimedia Commons.
I had uploaded photographs, organized Wiki Science competitions, and knew the basics of licensing and file uploads.
The Commons sessions completely changed that perspective. Rather than focusing on uploading more images, the discussions emphasized documenting knowledge in ways that remain useful for future contributors. We explored why metadata, categories, descriptions, geolocation, and licensing all play an essential role in making media discoverable and reusable. One idea particularly stayed with me. Commons does not necessarily need another photograph of a monument that has already been documented hundreds of times. It needs photographs that document the subject well.
That simple idea changed the questions I ask before uploading an image. Instead of asking whether I can upload a photograph, I now ask whether it genuinely helps someone understand a place, object, or tradition better.
The photo walk around the IIIT Hyderabad campus gave participants an opportunity to apply these ideas immediately. Rather than simply taking attractive photographs, we practiced documenting subjects from angles that communicated information clearly and added educational value.
The licensing session also helped demystify Creative Commons licenses through practical examples, making it easier to understand how open licensing enables collaboration across Wikimedia projects.
Good communication keeps communities growing
The final day focused on communication an area that is often overlooked but essential for sustaining volunteer communities.
A session on Visual Storytelling in Practice demonstrated how photographs and personal stories can help communicate knowledge more effectively than facts alone. It reminded me that contributors often remember stories long after they forget presentations.
Another workshop explored communication strategies for different audiences. Working in groups, participants designed outreach plans tailored to specific communities rather than relying on a single approach for everyone. That exercise reinforced a simple but important lesson:
There is no single Wikimedia audience. Students, teachers, heritage enthusiasts, language learners, and professionals all engage with Wikimedia for different reasons. Effective outreach begins by understanding those motivations.
The session on Wikivoyage also introduced me to another Wikimedia project that I had not previously explored in depth. It demonstrated how documenting travel knowledge, local culture, and places contributes to the broader free knowledge ecosystem.
Finally, a session on community communication platforms highlighted how mailing lists, Telegram groups, discussion forums, and social media help communities stay connected long after events have ended.
Building a community is not only about organizing an event. It is about creating ongoing conversations.
Learning by doing
One aspect of TTT that I particularly appreciated was its emphasis on practical learning.
Rather than relying entirely on presentations, participants engaged in group discussions, storytelling exercises, communication planning, case studies, and hands-on Commons activities.
These exercises encouraged us to apply ideas immediately instead of simply listening to them.
The collaborative nature of the program also created opportunities to learn from participants representing different language communities, projects, and experiences across the Wikimedia movement.
That diversity of perspectives became one of the most valuable parts of the training itself.
What changed after TTT?
The impact of Train the Trainer extended well beyond the three days of the program.
It changed how I think about community building.
Instead of focusing primarily on organizing events, I now think more about contributor retention, mentorship, and creating welcoming spaces for newcomers.
It changed how I contribute to Wikimedia Commons.
I now pay much greater attention to documentation quality, metadata, licensing, and preserving local heritage through meaningful photographs.
It also influenced my later work within the Wikimedia movement. Many ideas that I developed while preparing proposals for WikiConference India, as well as my growing interest in documenting local heritage and strengthening community communication, can be traced back to discussions and activities during TTT.
Most importantly, the program encouraged me to think beyond individual edits and toward strengthening the communities that make those edits possible.
Why programs like Train the Trainer matter
Every Wikimedia community faces different challenges, but many of those challenges share common themes: welcoming newcomers, retaining contributors, building trust, documenting knowledge responsibly, and communicating effectively.
Train the Trainer creates a space where volunteers can exchange experiences, learn from one another, and return home with practical ideas that can be adapted to their own communities.
For me, the greatest takeaway was not a single workshop or activity.
It was a shift in perspective.
Wikimedia is sustained not only by articles, photographs, or software, but by people who collaborate, mentor, listen, and continue learning together.
That is what I brought home from Train the Trainer 2026 and it is why I believe programs like TTT continue to play an important role in strengthening the Wikimedia movement.
Here is a quick overview of highlights from the Wikimedia Foundation since our last issue on July 3. Previous editions of this bulletin are on Meta. Let foundationbulletin@wikimedia.org know if you have any feedback or suggestions for improvement!
Wikimania 2026: Wikimania is happening this week! After the event, all streamed sessions will be linked in the program on Eventyay and later uploaded to Commons.
Grantmaking: The Global Resource Distribution Committee has published a Grantmaking Strategy draft that sets out a renewed approach to how the Wikimedia Foundation’s Community Fund is distributed across the Wikimedia Movement. The GRDC is requesting feedback from volunteers and affiliates, regardless of whether they are grantees or not.
Movement Ecosystem: A proposal that would update movement affiliate recognition and establish new, connected criteria for eligibility to receive Community Fund grants is now available for community review. You are invited to read the proposal and participate in the discussion until August 7.
Women+ contributions in Wikimedia Tech: A guide based on lived experience on how to address some of the invisible barriers for women+ in more technical Wikimedia spaces and recommendations to become more inclusive. Help further by filling out this survey to better understand technical contributions by women+ across Wikimedia projects until July 20.
Structured Experimentation: A reflection on the first year of structured experimentation highlights successful experiments such as Paste Check, Reference Check, and Tone Check, which improved editing outcomes and have been rolled out to more users, as well as experiments that did not lead to product changes.
Revise Tone test ended: The A/B test of Revise Tone ended on July 9. It showed that newcomer task completion rates increased by 38.7% compared to the default Copyedit task, with no decrease in edit quality. The feature is now available for everyone on the Arabic, English, French, and Portuguese Wikipedias. The plan is to release Revise Tone to more wikis.
Wikidata: The latest Wikidata Platform newsletter (July edition) shares how to identify and rewrite queries that rely on Blazegraph-specific extensions and affected by the migration off Blazegraph.
Discussion Tools: On English Wikipedia, DiscussionTools‘ Usability Improvements has now become default for talk pages. You can opt-out of these changes at any time in user preferences. With this, Discussion Tools are now fully available at all wikis.
Tech News: The latest highlights from Tech News week 28 and 29 include the new Parsoid parser continues to be deployed to additional wikis, making it easier to introduce new reading and editing features. See also the 72 community submitted tasks that were resolved over the last two weeks. Overall, from April – June 2026 about 337 community tasks were resolved by the Wikimedia Foundation.
Language inclusivity at Wikimania 2026: New approaches to translation and interpretation will be tested at Wikimania this year.
Call for submissions open for Wikimedia Latin America Conference 2026: The Wikimedia community in Latin America has opened the call for session proposals for the Wikimedia Latin America Conference 2026. Community members are invited to submit proposals for the conference program by August 10.
UK Online Safety Act: Ofcom, the United Kingdom’s Office of Communications, announced that Wikipedia is not designated as a Category 1 service under the Online Safety Act (OSA). This is an important and welcomed outcome as a Category 1 designation could have included privacy and safety risks to our global community of volunteers.
UN Open Source Week edit-a-thon: Volunteers created 60 new Wikipedia articles and made nearly 700 updates to improve Wikipedia’s coverage of UN and open source topics at the second UN Open Source Week edit-a-thon co-hosted by Wikimedia Foundation.
“Don’t Blink”: The latest developments from around the world about protecting the Wikimedia model, its people and its values.
“A Wiki Minute” videos: New videos are added to the series answering some of the most common questions such as “Do you still need Wikipedia when AI can answer anything?” and “Does Wikipedia push a political agenda?”.
Wikipedia 25 brand collaboration in Indonesia: On 4 July, the Jakarta-based street wear company Ageless Galaxy launched a Wikipedia 25 collection, the first ever Wikipedia Brand Collaboration in Asia. The apparel collection featured hats, t-shirts, and a jigsaw puzzle cardigan. Wikimedians in Indonesia joined Ageless Galaxy for a launch party.
Board Elections Eligibility: There are new proposed eligibility criteria for standing as a candidate in Wikimedia Foundation Board of Trustees election now available for feedback. The requirements are more specific and detailed than in years past, to both inform the community of what the Board needs and to create multiple pathways to the Board for Wikimedians.
This conference is a space to create, share, and lead, not only to attend. This year, we are building around the theme Re-imagining the Knowledge Commons, and extending an invitation to you to find out together how we build and care for shared knowledge. We’re looking for proposals that go further than traditional editing, uploading, and coding, trying out new ways of taking part and making knowledge as a community.
WikiConference India will happen in Kochi, Kerala on 4th- 6th September 2026, in the spirit of Namukku Othukoodam (Let’s meet!).
The Evolving Knowledge Commons
The Knowledge Commons, in its digital form, has come to be defined by the Wikimedia ecosystem. Emerging trends in AI-based search and summarization, content generation, language translation, and new forms of interaction have brought new challenges to the way in which we contribute to the knowledge ecosystem and the way in which consumers unknowingly draw from it.
To ensure that knowledge remains open, inclusive, safe, and accessible, we cannot simply maintain what exists: we must actively shape ourselves to tackle what comes next and this is at the core of our Programming vision. Our theme aligns closely with broader Wikimedia movement priorities and also the global trends tracked by the Wikimedia Foundation. In a time of declining trust in online information and the rapid rise of AI-generated content; strengthening community-led, human-created knowledge becomes more critical than ever.
Key Questions We Are Exploring
How efficiently are we bringing traditional stores of knowledge- from people, libraries, academia, galleries, and museums- into the Knowledge Commons?
How do we protect the integrity of our knowledge, keep the Knowledge Commons verifiable, and protect the agency of our contributors?
What do all these changes and future possibilities mean for those of us contributing to Wikimedia projects focussing on Indic languages?
Core Focus Themes
We welcome submissions that address our core objectives and align with the following strategic directions:
Strategic Roadmap Building: Sessions that help us navigate the shifting digital landscape, the rise of new technologies, along with Sustainability and future pathways for Wikimedia in India and South Asia.
Community Leadership and Governance: Sessions that highlight community-led initiatives and how grassroots leadership can sustain open knowledge.
New Models of Participation and Contribution: Sessions related to technological enhancement, GLAM, EduWiki, and Knowledge-resource-content partnerships.
Knowledge Equity and Inclusion across Languages and Regions: Sessions focused on strengthening Indic-language projects and ensuring underrepresented histories remain at our heart.
Technology, Tools, and the Evolving Digital Knowledge Ecosystem: Exploring infrastructure, platform engineering, and the role of AI.
We especially encourage proposals that:
Share collective practical experiences and learnings, offering generalizable knowledge.
Highlight independent/community-led innovations that are reproducible in wider contexts.
Create space for dialogue, skill-building, and co-creation that travels beyond a single community.
If you have received a scholarship to attend, we especially encourage you to propose a session or be a part of one! Whether you are an experienced contributor or a newer community member, your ideas and experiences are essential to building a meaningful and inclusive program.
Submission Tracks & Formats
To help organize our collective program, we invite proposals across a variety of formats. Please review the session formats and suggested durations below to see where your idea fits best:
Poster presentations: Posters will be presented at the Conference Venue, allowing one-on-one discussions with participants.
Community meetups
Note: Unconference sessions are intended for informal meetings after 5:00 PM (following main conference hours) and will not be listed in the primary program schedule. Even if your idea does not fit perfectly into one of these exact buckets, we still want to hear it!
“WikiConference India is, at its core, a space for communities to meet, reflect, and reimagine- not just a conference, but a moment where communities see themselves and the impact of their work more clearly. WCI 2023 brought that vision to life by bringing Wikimedians from across India and South Asia together in Hyderabad for the first time since 2016, under the theme ‘Strengthening Bonds.‘ It reminded us that reimagining the knowledge commons is ultimately about people, how we connect, collaborate, and care for the spaces we build together. We hope WCI 2026 continues this journey forward.” – Nitesh and Nivas, Organisers of WikiConference India 2023
“WikiConference India 2026 is shaped by the community, and the program is at its heart. Through this call, we invite contributors to bring their ideas, experiences, and questions to the table and help co-create the conversations that matter most.”
As we prepare to meet in Kochi this September, remember that the “Knowledge Commons” belongs to you. We cannot reimagine it without your voice, your leadership, and your creativity. Submit your proposal today and help shape the roadmap.
When I first came across the call for applications for the Wiki Afrodemics Mentorship Programme, I almost scrolled past it. I had seen calls like this before, applied to a few, and heard nothing back. But something about this one felt different maybe it was the clarity of the focus countries, or maybe it was simply timing. I applied anyway, not expecting much, and a few weeks later found my name on the list of Cohort 1 fellows for the Anglophone group.
That single email changed the next three months of my life on Wikimedia.
The Application and the Wait
Applying was straightforward enough: a form, a short statement of interest, a bit about prior contributions. The harder part was the waiting. Mentorship programmes like this one are competitive, and I remember refreshing my inbox more often than I’d like to admit. When the selection list finally came out, and I saw my name among the twenty participants chosen from across the five focus countries, it felt like validation not just of my interest in Wikimedia, but of the small, scattered edits I had been making before anyone was watching.
A Setback I Didn’t See Coming
Not long into the programme, I ran into a wall I genuinely didn’t expect: I was temporarily blocked on Wikimedia. I won’t pretend that moment didn’t sting. For a brief period, I couldn’t edit directly, and I had to sit with the uncomfortable feeling of being sidelined from the very project I had just been selected to contribute more to.
But this is where the structure of the programme and the patience of the mentors made all the difference. Instead of treating the block as a dead end, I learned to treat it as a detour. I kept writing. I drafted articles offline and in my sandbox, and submitted them for review by experienced Wikimedians who could push them through on my behalf or guide me on how to get back in good standing. It taught me something I didn’t expect to learn from a mentorship programme: that contributing to Wikimedia isn’t only about the edit button. It’s about the work itself: the research, the sourcing, the drafting, and there is always a way to keep that work moving, even when the front door is temporarily closed.
Mentors Who Knew How to Teach, Not Just Tell
If there’s one thing that stood out across the three months, it’s the quality of the mentorship itself. It’s one thing to know Wikimedia inside and out; it’s another to be able to teach it well. Our mentors managed both.
The Wikidata sessions, in particular, stretched my thinking. Wikidata isn’t always intuitive when you’re coming from a Wikipedia-first mindset, but the structured, month-long training broke it down in a way that made the structured data side of the movement click for me. Moving between Wikipedia, Wikidata, Wikimedia diff and Wikimedia Commons over the course of the programme gave me a much fuller picture of how the projects connect, how an image uploaded to Commons, a claim modeled on Wikidata, and an article written on Wikipedia all reinforce each other.
Ibjaja055, CC0, via Wikimedia CommonsIbjaja055, CC0, via Wikimedia CommonsIbjaja055, CC0, via Wikimedia CommonsIbjaja055, CC0, via Wikimedia Commons
The Lesson I’m Taking With Me: Notability
Of everything covered across the cohort – content gaps, gender representation, cross-regional collaboration – the single most valuable thing I walked away with is a real, working understanding of Notability.
Before this programme, notability was a word I associated with rejection, articles I have seen tagged for deletion, drafts that sat untouched because I wasn’t sure they would survive scrutiny. Now I understand it as a framework rather than a gatekeeping obstacle. Knowing how to evaluate whether a subject meets Wikipedia’s notability guidelines, and how that maps differently onto Wikidata’s notability expectations, has changed how I choose what to write about in the first place. I no longer draft and hope. I check first, source deliberately, and build a stronger case for inclusion from the very first sentence.
That shift alone from guessing to evaluating is worth everything else I gained from this cohort.
Looking Ahead
As Cohort 1 closes out for 2026, I am left with a mix of gratitude and momentum. Gratitude for mentors who gave their time generously over three months of training, listening, and patient correction. Momentum because I now have both the skills and the confidence to keep contributing, blocks, setbacks, and all.
I am already looking forward to the 2027 cohort, not as a participant this time, perhaps, but as someone who might be in a position to give back the same kind of guidance I received.
For over ten years, Wikimedia communities and associations in Europe have been working together to improve the legal framework for free knowledge in Europe – advocacy at the European level and participation in shaping policy at the national level go hand in hand. Copyright reform, the AI Act, and the Digital Services Act (DSA) are just a few examples that have occupied us in recent years.
Claudia Garád, Wikimedia Österreich’s Executive Director, is an active supporter of the European idea within the Wikiverse. Since 2022, Claudia has served as the volunteer President of Wikimedia Europe – a European umbrella organisation in Brussels, shaped and managed by over 30 European Wikimedia associations, and acting as a strong voice for public-interest-oriented internet policy at the European level.
Claudia, what exactly does Wikimedia Europe do?
Our aim is to be a kind of network hub in the international Wikiverse, facilitating effective exchange and cooperation between Wikimedia organizations. Decentralized cooperation is a major strength of our volunteer communities. In Europe, we demonstrate that Wikimedia organizations also have this way of working in their DNA. Especially in these politically and socially challenging times, when civic engagement is increasingly restricted everywhere, we can only survive and exert real influence as a collective.
Wikimedia Europe is therefore a shared platform for collaborative work on European legislative processes that impact Wikimedia projects. It also supports collaborative fundraising activities for joint projects and skills development and transfer. This happens particularly in the areas of advocacy and fundraising, but also beyond.
Why is international digital policy important for an affiliate such as Wikimedia Austria?
First and foremost, it’s about the digital infrastructure on which everything is built: For Wikipedia, other Wikimedia projects, or our ÖsterreichWiki to function at all, we need an open, free, and globally functioning internet. Without interoperability and open standards, these projects would not be possible. European and international digital policy, therefore, determines very concretely whether free knowledge remains accessible to everyone and under what conditions our communities can operate.
Furthermore, national legislation in this area is largely based on European regulations, which are usually designed with large, commercial platforms and social media in mind. Non-commercial, public service digital projects are often not given a voice and cannot lobby to the same extent as multi-billion-dollar corporations. Through Wikimedia Europe, we have the opportunity to pool and amplify the activities and resources of individual members and thus counteract this imbalance.
Last but not least, it’s also about transparent political decision-making processes: Digital policy doesn’t just affect technical or economic developments, but directly impacts how we live together as a society. After all, we use digital platforms for consumption and communication, we form our opinions online, we work with digital tools, and so on. Therefore, we advocate for the structural integration of organizations representing civil society into decision-making processes in Austria and Europe: Politics, business, science, and civil society should work together transparently on solutions.
What is your role at Wikimedia Europe?
During the founding phase, I initially served as president on the interim board to prepare the organization for its spin-off, and then in 2025 I was elected to the first “official” board in the same role. My responsibilities include representing the association externally—although this has now been largely delegated to the managing director, Anna Mazgal, in day-to-day operations. In addition, I handle governance and HR matters, as well as risk management within the association, and act as a liaison to the global Wikimedia movement beyond Europe. Last year, together with the board, employees and a project group of representatives from the member organizations, we also developed our first integrated multi-year strategy.
What does Wikimedia Europe mean to you personally?
Jacques Delors, former President of the European Commission and architect of European integration, used to say: “Never choose between being an optimist or a pessimist. The only choice you can make is to be an activist.”
My volunteer role as President of Wikimedia Europe gives me the opportunity to be an activist in two projects that have been formative for me personally and represent, in my opinion, the most wonderful experiments in human history: Wikipedia and the European Union. Both are examples of how people can achieve the previously unimaginable when they consistently prioritize trust and cooperation. This gives me courage, confidence, and a sense of self-efficacy, even on dark days when I feel the world is increasingly falling apart.
To learn more about WMEU visit our page on Meta-wiki
Check the WMEU public policy blog to follow the latest developments in our public policy work in Europe
Peter Zlabinger is a Communications Advisor at Wikimedia Österreich and his job is to explain the ins and outs of the Wikiverse to our various audiences and stakeholder groups. In EU advocacy two complex universes meet – the policy making of the European Union and the Wikimedia Movement. Wikimedia Österreich has been an integral part of Wikimedia advocacy in Europe. In this interview Peter explores what makes advocacy so exciting but also important from an Austrain but also global movement perspective.
The Wikimedia movement’s commitment to preserving and sharing knowledge took on a uniquely South African flavour last month as newly appointed Wikimedia Foundation CEO Bernadette Meehan embarked on one of her first major international visits since assuming the role.
Accompanying her was Bobby Shabangu, a respected South African open-knowledge advocate and the first African ever elected to the Wikimedia Foundation Board of Trustees. Together, they spent time with Wikimedia South Africa volunteers, partners, and community members, exploring how local initiatives are helping ensure African languages, histories, and cultures are represented online for generations to come.
Cape Town: Where Ancient Stories Meet the Digital Future
The South African tour began in Cape Town with a visit to the Iziko Museum, where participants experienced a guided tour that perfectly reflected the Wikimedia movement’s belief that knowledge is created, shared, and preserved by people.
Among the highlights was a presentation of some of South Africa’s oldest rock art. Created over centuries by multiple generations, these artworks offered a striking parallel to the way Wikipedia itself is built today — collaboratively, incrementally, and collectively. Just as countless hands contributed to preserving stories on stone, volunteers across the world continue to build humanity’s largest collection of freely accessible knowledge.
Conversations with Cape Town-based Wikimedia South Africa board members and partners took visitors further “down the WikiRabbit Hole,” exploring innovative language preservation projects currently underway. Particular interest was shown in efforts supporting the incubation of Afrikaaps, a Cape Town dialect of Afrikaans, as well as initiatives enabling KhoeKhoe language speakers to contribute directly online through the development of specialised keyboards and editing tools.
Johannesburg: Showcasing Local Innovation
The visit continued in Johannesburg, where Wikimedia South Africa members presented a range of projects addressing the digital knowledge gap facing African languages and communities.
Many South African languages have rich oral traditions but limited representation in written and digital archives. Community members demonstrated how Wikimedia projects are helping address this challenge through locally led initiatives.
Decolonising Language Through Knowledge Creation
Editors showcased work on smaller and underrepresented language editions of Wikipedia, including Siswati, Tshivenda, Sepulana, and Xitsonga. Through the creation of high-quality educational content in local languages, volunteers are helping ensure that knowledge is accessible to communities in the languages they speak and understand.
Jo’burgpediA and Community Partnerships
The chapter also highlighted collaborations with universities, libraries, museums, and cultural institutions. Through initiatives such as Jo’burgpediA, students and community members are trained to document local history, ensuring that communities can tell their own stories rather than having them told by others.
Preserving Oral Knowledge
Discussions also focused on innovative approaches to documenting verified oral histories and indigenous knowledge. These efforts challenge traditional archival models that have often excluded African perspectives and lived experiences, creating new pathways for preserving cultural heritage within Wikimedia platforms.
A Global Mission Powered by Local Communities
Addressing chapter members, Bernadette Meehan emphasised the Foundation’s commitment to community-led knowledge creation and governance:
“Wikimedia’s global mission relies entirely on the local communities who understand the nuances of their own culture. Listening to the brilliant initiatives run by Wikimedia South Africa proves that the future of free knowledge is inherently multilingual and diverse.”
Her remarks recognised the vital role that local volunteers play in ensuring the internet reflects the richness and diversity of the communities it serves.
For Bobby Shabangu, the visit was also a reminder of how far the African Wikimedia movement has come.
Reflecting on his own journey, which began with editing Siswati Wikipedia because, as he puts it, “if I didn’t edit it, no one would,” Shabangu highlighted the growing influence of African contributors within global internet governance and knowledge-sharing spaces:
“We are moving past the era where Africa is merely a consumer of global knowledge. Through the hard work of our chapter members, we are ensuring that our grandfathers’ stories, our traditional customs, and our beautifully diverse languages are permanently archived for the next generation.”
Looking Ahead
The visit reaffirmed the importance of community-driven knowledge creation and the growing role South Africa is playing in shaping the future of free knowledge globally.
From ancient rock art in Cape Town to digital language preservation projects in Johannesburg, the message was clear: preserving knowledge is not only about recording the past — it is about ensuring that future generations can access, contribute to, and share their own stories.
As Wikimedia South Africa continues to champion linguistic diversity, cultural heritage, and open knowledge, partnerships between local communities and the Wikimedia Foundation will help ensure that South Africa’s rich tapestry of languages and cultures remains visible, accessible, and thriving online for decades to come.
Wiki Indaba 2026: Passing the Baobab
Joining the Wikimedia South Africa chapter were members of the core organising team for Wiki Indaba 2026 from Côte d’Ivoire, including Emmanuel Gueh and Donotein. Their participation provided an opportunity not only to engage in discussions with the South African community but also to share an update on preparations for Wiki Indaba 2026, Africa’s premier gathering of Wikimedians.
A special highlight of the visit was the ceremonial handover of the Wiki Indaba Baobab Tree statue to the new organising team. The Baobab, often referred to as the “Tree of Life” in Africa, symbolises wisdom, resilience, community, and the sharing of knowledge—values that closely align with the spirit of Wiki Indaba and the Wikimedia movement.
Wikimedia South Africa leadership proudly passed the statue to the Côte d’Ivoire organising team, marking the beginning of what is hoped will become a lasting Wiki Indaba tradition.
As the creator of Wiki Indaba, Wikimedia South Africa Board Chair and long-time Wikimedian Dumisani Ndubane reflected on the significance of the moment. He expressed his pride in inaugurating this new tradition and shared his hope that, in years to come, the Baobab statue will continue its journey across the continent before eventually returning to South Africa when Wikimedia South Africa once again hosts Wiki Indaba.
The handover symbolised more than a transfer of responsibility—it represented the continued growth of the African Wikimedia movement, the sharing of leadership across communities, and a collective commitment to ensuring that Africa’s knowledge, languages, cultures, and stories remain visible and accessible to future generations.
I love Wikipedia, command-line interface (cli) and respect offline stuff. Finally I did it – Rust software (for peformance), code generated by llm gpt-5.5 xhigh.
How to use: download the Wikipedia dump, and
cargo run --release -- path/to/dump.xml -o path/to/output_directory
Yes for offline reading we have Kiwix – but terminal is also good. For you batery as well. Or when you system in under heavy load so your CPU/RAM are limited. Man format mean that you can read Wikipedia through SSH as well, and from very cheap devices.
This is not perfect – we should add special handlers for different templates. But it already mostly works and valuable for the comunity. Try it, ping if you want to improe something. You are welcome to pack this software for you linux distribution – I packed only for Gentoo.
This is a testimonial graphic design for the On-Wiki Skills Mentorship Program Cohort 2 (2026)
Before joining the On‑Wiki Skills mentorship program, my engagement with Wikimedia projects was mostly as a reader. I was curious about open knowledge but unfamiliar with how data works or how contributors actively build and maintain Wikimedia projects.
Through the mentorship, I was introduced to Wikimedia as a collaborative ecosystem that offers structured ways of organizing and connecting knowledge. The learning curve was initially challenging especially understanding items, properties, statements, and references but with guidance from mentors and hands‑on practice, these concepts gradually became clearer.
Skill Acquired
During the training, I developed skills in sourcing and referencing information and understanding structured data. These experiences changed how I view knowledge creation and reinforced the idea that learning continues beyond formal lessons.
Some of my image uploads
Administration Block from Ghana Standards Authority
Pesticides Residue Block from Ghana Standards Authority
Fish Department from Ghana Standards Authority
Participating in practical exercises and guided contributions helped build my confidence and showed me that I could engage meaningfully with a global open‑knowledge communication.
My Gratitude
I am grateful to the Wikimedia community, mentors, and program organizers for their support and encouragement. Completing the mentorship marked a new beginning for me, and I look forward to continuing to explore and contribute to open knowledge
For years, rites and rituals have shaped the identity of communities across Nigeria. From naming ceremonies and traditional weddings to religious observances and everyday customs, these practices preserve values, beliefs, and collective memories. Yet many of these cultural expressions remain underrepresented on Wikimedia projects.
Wiki Loves Africa 2026 for Creatives in selected Nigerian universities broughtt together photographers and storytellers to document and share Nigeria’s rich cultural heritage under the theme “Rites and Rituals.” The campaign was held from February to April 2026.
Creating awareness and recruiting creatives
Between 1 and 28 February 2026, the campaign focused on awareness and participant recruitment. Outreach activities targeted creatives across six institutions and communities, while social media campaigns and sponsored advertisements on Facebook and Instagram helped expand participation.
These efforts introduced creatives to Wikimedia Commons and encouraged them to contribute freely licensed media that celebrate Nigeria’s diverse traditions.
Flyer for Awareness
Launching the campaign and introducing Participants to Wikimedia Commons
The campaign officially kicked off with a virtual launch held on 4 March 2026. The session introduced participants to the Wiki Loves Africa 2026 theme, “Rites and Rituals,” and provided an orientation on contributing to Wikimedia Commons.
Participants learned about:
Wikimedia Commons and its role in preserving open knowledge;
Creating Wikimedia accounts;
Understanding free licenses;
Best practices for successful contest submissions.
The virtual launch provided a foundation for participants, which some were first-time contributors to Wikimedia projects.
Training sessions
Throughout March and April, participants joined several training sessions organized by the international Wiki Loves Africa team. These sessions equipped creatives with technical and ethical skills required for documenting cultural heritage responsibly.
Audio creation and sound design training for Wikimedia Commons on 19 March 2026;
Ethical storytelling training on 27 March 2026;
“Documenting our Rites and Rituals with Respect and Integrity” on 28 March 2026;
Pattypan mass upload training on 4 April 2026;
Uploading photographs, videos, and audio files to Wikimedia Commons on 10 April 2026.
These learning opportunities strengthened participants’ understanding of visual storytelling, copyright, ethical documentation, and the technical processes involved in contributing multimedia content to Wikimedia Commons.
Taking the campaign to communities through photowalks
Beyond virtual engagement, the campaign emphasized practical documentation through physical meetups and photowalks.
Creatives in Rivers State held their meetup and photowalk between 3–5 April 2026, where participants explored their communities and documented the Easter Rites.
On 11 April 2026, creatives in Abuja gathered for their meetup and photowalk, providing opportunities for collaboration, peer learning, and hands-on experience in capturing images related to the theme. Gombe creatives hosted their physical meet-up on 25th April 2026, and so did Edo State and Oyo State.
Physical meet-up of creatives in AbujaDuring physical meet-up
Every photograph, audio recording, and video uploaded to Wikimedia Commons contributes to a growing repository of freely accessible knowledge, ensuring that Nigeria’s rites and rituals are visible to people around the world and preserved for future generations.
As participants continue contributing beyond the campaign, Wiki Loves Africa remains a reminder that documenting culture is not only about preserving the past. It is also about ensuring that African stories are represented and shared openly with the world.
Commons Quality Images (COM:QI) was created during June 2006. The purpose of COM:QI was to identify high quality photographs taken by Wikimedians and made freely available to the community at large through Commons. As of 14 June 2026 over 450,000 media files, mostly photographs, have been recognised through this process. Alongside photographs, QI has Microscopic images, animated GIF’s, and graphical works including our own QI seal (pictured).
Commons Quality Image seal
From the Beginning
In September 2005 Commons became live, just 6 months later I was encouraged over by User:Pfctdayelise who had been there from the first month. In June 2006 User:Pfctdayelise was searching to put together collections of images created by Commons users as rotating background images, or calendars to help promote the project. I tried to help, we could not even find a common subject matter consisting of reasonably sized and quality printable images. Wikimedia Commons had some great photos, some of which had been featured, but a substantial portion were images scraped from other sites like NASA and Flickr.
That started me thinking about how we raise the profile of community uploaded images and encourage efforts to improve them. Having already contributed to some Good Articles over on en.Wikipedia I thought we could create a similar project but for photographs.
Creating the project I saw some key necessities: it needed to be efficient with clear standards, while not taking weeks to decide the outcome. Most importantly it needed to focus only on images that had been taken by members of the Commons Community and that once recognised it could not be taken away. From there working with User:Wikimol and others we built the project. We also created image guidelines which would enable consistent reviewing of nominations.
Low resolution / thumbnail. Photographic QP has to have at least 1.92 megapixels (=1600×1200, for example)
JPEG problems. Too much compressed, too low JPEG quality settings in camera / when saving. Visible jpeg artifacts. => Use better quality settings (e.g. set JPEG “superfine”, shoot RAW, save in photoshop with max. quality)
Noise problems, too much noise. Be it chroma noise, luminance noise, visible grain, scratches in scans… QP should not have distracting amount of noise when viewed in 100%
Bad exposure. Overexposure, blown out highlights, underexposure, shadows details replaced by jpeg maps… In incorrectly exposed images, significant details in a significant part are lost.
Color problems. Bad white balance. Distracting (typically purple) hazing at 100%. Color aberration. QP must have reasonable colors (which does not necessarily mean natural colors).
Improper or undefined focus, insufficient depth of field. QP should have clearly defined focus, e.g. main subject in focus, foreground and background out of focus. Or the whole scene in focus. Counterexample – main subject blurry, foreground even more blurry, focus is somewhere between main subject and background. DOF could be low on purpose.
Blur. Images blurred just because of shaking hand or subject moving too fast. Motion blur in QP has to have purpose.
Poor lighting. Including: distracting reflections (usual problem with built-in flash), unintended vignetting, distracting harsh shadows. Generally bad lighting makes scenes with space look flat.
Overfiltered. There are so many PS/Gimp filters. Rarely a better image is created just by applying more and more filters…
Bad or nonexistent composition, unclear or nonexistent subject. QP should have subject and composition of the image should support depiction of that subject, not distract from it.
Bad perspective, tilt, and other distortions. An eye (or, more precisely, a brain) is a sensitive detector capable of spotting even a small tilt … falling trees, churches, inclined water surfaces,… Images of architecture should usually be rectilinear and without too much perspective distortion.
Stitched images, panoramas.
Panoramatic QP has to have a height of 800px min.
Stitching problems. Stitched images should be without artifacts, colors and lightness should be the same across the image.
All of these may look familiar to many of you as they are still present as Commons:Image guidelines, the very same guidelines that every WikiLoves or similar competition uses as their rules.
Now we had a process with guides, we approached User:LadyofHats who was known for the drawings she was uploading at that time. LOH was asked to create a seal which we could use to help identify successful photos, that seal is the one we continue to use. Some time during this the QI seal itself got recognised as a Quality Image.
In those early months every image was reviewed and the pages manually updated, a time consuming process that would occasionally be interspersed with edit conflicts. The community embraced QI and a new project called Valued Image emerged for recognising sets of image rather than individual images. It was the efforts of User:Dschwen, who had also taken on maintenance tasks, decided to create a bot that would do most of the work for us. Mike Peel continues to maintain the QICbot and has been invaluable at keeping it working, as of 16 June 2026 QICbot has performed 922,000 edits looking after QI tasks.
QI grew fast once the QICbot came on board. I stepped into the background watching the community grow QI. Over time many contributors have made QI. Like User:Poco a poco who presented at Wikimania in London on how QI helped him improve his contributions. For the curious there is an opt-in unaudited list of QI by photographers. Whether it’s one successful image or 20,000 of them they all make QI what it is.
The future
I think some of QI’s potential still hasn’t been realised. I saw it as a historical record of the growth of photography, of something researchers could look back over and see how it has matured. Perhaps more could be done to integrate QI images in content on other projects as they are recognised as our better works. Maybe a bot could triage nominations to identify the more regular issues that cause images to be rejected, though the human touch should always remain the final judge.
In the future, QICbot will reach 1 million edits, QI will reach 500,000 very soon. QI carries within itself a lot of untapped potential.
Could it be time to consider creating a sister project for Quality Videos? I know one thing: there will always be enough high quality images created by Wikimedians to make a calendar for any subject. The challenge will be in choosing just 12.
I would must take a moment to acknowledge everyone who has already helped Commons Quality Images along it’s 20 year journey there has been so many as well as everyone who joins the efforts in future. It’ll always be nice to reflect and be able to say I was able to open the door to something so special. The future of Quality Images is now a journey the Wikimedia Commons Community will decide.
QI’s prosperity comes from the collective effort of everyone! Perhaps this will encourage you to join the QI club.
For 25 years, the Wikimedia movement has been constantly innovating to promote the values of free and universal sharing of all human knowledge for the benefit of humanity. In 2026, as we face the challenges of connectivity, misinformation, and the rise of AI, our ability to innovate collectively is more essential than ever.
To address these developments, we are introducing a new program focusing on Team Challenges in the lead-up to Wikimania. This year, we are inviting experts and newcomers to the Wikimedia community to join forces with participants in our Wikimania Hackathon. The idea is not to replace the hackathon format, but to offer a different approach. We want to bring together diverse perspectives and skills by pairing Wikimedians from all walks of life (not just developers!) with professionals from other fields. The goal is to share our best practices and design digital tools that are sustainable, inclusive, and accessible to all.
Next week, we are kicking off our first online orientation sessions training newcomers on wiki tools. In the coming weeks, they will form teams with Wikimedia community members to undertake one of the 2026 challenges. Stay tuned for further updates as the teams work together toward the Team Challenges showcase concurrent with Wikimania Paris.
We can’t wait to see what the cross-pollination between different disciplines, perspectives, and skill sets will bring. Good luck to all!
10 challenges for maximum impact
The goal is simple: To turn the community’s needs into concrete solutions. We have selected 10 major technical challenges that address the complex issues encountered on a daily basis across Wikimedia projects.
These challenges are based on the three pillars of Wikimania 2026:
FREEDOM: Ensuring the interoperability of free software to facilitate its reuse and maintenance, and to guarantee its long-term viability.
EQUITY: Promoting multilingualism. Designing tools capable of overcoming language barriers to bridge gaps in sources and contributions on a global scale.
RELIABILITY: Developing AI-based detection tools to assist volunteer moderators in their work of verifying and protecting information.
Discover the 10 Challenges
Fix the sources Automatically detect outdated, retracted or broken sources in articles, and facilitate their updating or archiving.
Welcoming newcomers Reinvent how we onboard newcomers through chatbots, guided micro-edits, and gamification lower the barrier to entry on wikis.
Update the obsolete Identify and flag outdated articles, obsolete illustrations and screenshots, and outdated graphics to: encourage the community to update them.
Boost multimedia experience Enhance the audiovisual experience on wikis: playlists on Commons, a modernised video player, annotation of audio segments, and games focused on accents and pronunciations.
Deciphering biases Develop tools to detect gender, geographic, and cultural biases in Wikimedia content to: encourage the community to correct them.
Gamify knowledge Create serious games and playful experiences using Wikimedia data to: learn, contribute, and explore in new ways.
Stream data with Wikidata Take advantage of Wikidata content to: generate lists, tables or up-to-date visualisations, semi-automatically, for other projects in the Wikimedia galaxy or further.
The editor of the future Improve the editing experience: search & replace in the source editor, a standardized citation toolbar, a better visual table editor, and text alignment in VisualEditor to modernise and streamline contributions.
Explore knowledge Design and experiment new ways of sharing knowledge, in interactive and visual ways…for the modern web!
Connect multilingual knowledge Facilitate translation, the reuse of content across Wikimedia projects, and collaboration with other Open Source projects so that knowledge can flow freely in all languages.
For more information on the Team Challenges and how to participate, please visit the Team Challenges section of the Wikimania wiki. Please note that registration for the Team Challenges does not grant access to the rest of Wikimania.
Growth features are now available at Wikidata. This update enables access to Mentorship (if configured), Impact module, the Help Panel, and a simplified Newcomer Homepage (without Suggested Edits). Wikidata administrators are still configuring the features through Community Configuration.
Updates for editors
The special page Special:RangeCalculator has been created. It allows users to find an IP range without needing to rely on external tools. Until now, this tool was only available to CheckUsers. [1]
Sub-referencing is a new MediaWiki feature that allows editors to reuse references with different details. It will be deployed to most small and medium-sized Wikipedia language versions on June 23. The FAQ lists possible actions to take on your wiki to support the deployment. Check the rollout plan for the next deployment steps. [2]
Starting next week, users will get a notification when they are blocked or unblocked from editing, or if this block changes. [3]
Starting next week, abuse filters that are set to “require CAPTCHA verification” will begin to also affect users with the skipcaptcha right, which includes most autoconfirmed users. Bots are exempted. This change only affects edits that trigger an abuse filter. The skipcaptcha right will continue to exempt users from having to solve CAPTCHAs in the ordinary course of using the wikis. [4]
Reference documentation for the Lift Wing API has moved from the API Portal to the interactive REST Sandbox.
WikiConference North America 2026 will take place in Edmonton, Alberta (Canada), on September 25–27, 2026, with the Culture Crawl happening on September 24.
We expect around 250 people to attend. This will include scholarship recipients, guest speakers, and affiliates from across the region. The conference will feature workshops, presentations, and networking opportunities, fostering connections among Wikimedians, educators, and cultural organizations.
Building the Future of the Commons
“Building the Future of the Commons” is about reimagining how we create and share free knowledge in a rapidly changing world. As technologies like AI reshape how information is produced and consumed, the Wikimedia movement has a unique opportunity to ensure the commons stay open, human, reliable, and inclusive.
This theme aims to spark conversations across communities, the tech world, and cultural and educational ecosystems. Together, we will explore the evolving relationship between Wikimedia and AI, highlight underrepresented voices in our communities, and deepen our work with GLAMU institutions.
One of the conference days will also focus on empowering contributors at every level: from newcomers making their first edits, to everyday contributors looking to develop or improve their skills, to experienced users with extended rights and the issues that are most relevant to them.
Registration will open on Wednesday, June 3, 2026.
Italian wikimedians discussing web accessibility at the Wikimedia Hackathon 2026
Web accessibility is not merely a technical feature. It is a prerequisite for truly free knowledge. During the recent Wikimedia Hackathon 2026, held in Milan, we came together as a dedicated group hailing from Italy to confront a quiet yet persistent issue: the barriers that prevent visually-impaired individuals from fully engaging with Wikipedia and its sister projects.
Thus, Valcio, Daimona Eaytoy, and Piergiovanna Grossi (WMIT) led the unconference session “Wikipedia for Everyone: Closing the Accessibility Gap”, which served as both a wake-up call and a collaborative workshop. By examining how community-made templates and interface elements often fail our users, we aimed to transition from identifying problems to building sustainable solutions.
This is a short recap for those who missed it.
The Reality of the Digital Barrier
Home page for MediaWiki Accessibility Checker
The session opened with a candid look at the current state of our interfaces. While MediaWiki provides a robust foundation, years of community-driven customisation have inadvertently introduced many accessibility violations. Key issues discussed included:
Missing Alt-Text: Images essential for understanding content often lack descriptions or alternative text which is readable by screen readers, assistive technologies that read out graphic content to visually impaired users.
The “HTML Wall”: Many tables and templates lack proper semantic markup, forcing text-to-speech tools to read out raw code rather than structured information.
Contrast and Colour: Numerous gadgets and banners still fall short of the WCAG 2.2 AA (a web-accessibility standard) minimum contrast ratios, rendering them invisible to users with colour blindness or low vision.
Measuring Missing Alt-Text
The unconference session also sparked a small follow-up experiment. CristianCantoro set out to measure how widespread the issue of missing alt-text is on Italian Wikipedia and Lombard Wikipedia, combining the Wikipedia HTML dumps provided by Wikimedia Enterprise with the XML dumps published by the Wikimedia Foundation. The initial results confirm the scale of the challenge: more than 90% of images used in Italian and Lombard Wikipedia articles lack alternative text.
These numbers are a reminder that missing alt-text is still an open and large-scale challenge across languages. If we want Wikipedia to be truly open to everyone, we need better tools, workflows, and community practices to help editors add alt-text and meaningful descriptions to images.
From Discussion to Action: The MediaWiki Accessibility Checker
Logo for MediaWiki Accessibility Checker
To move from awareness to action, one of the session participants — Super nabla from the Indic MediaWiki Developers User Group — built a concrete solution during the hackathon itself. The tool, available on Toolforge, assists editors and developers in meeting accessibility standards: the MediaWiki Accessibility Checker. Try it out: https://accessibility-checker.toolforge.org/
Built on the industry-standard axe-core engine and Playwright, the tool is specifically adapted for the MediaWiki ecosystem. It allows editors and developers to (i) perform deep audits (queryable both from the frontend interface as well as from a dedicated RESTful API) based on WCAG 2.2 AA (and other standards) on any wiki URL, including project pages; (ii) generate professional reports in multiple formats, including PDF and Wikitext for easy sharing on-wiki; (iii) utilise a modern interface designed with the Wikimedia Codex design system, ensuring a seamless experience for contributors.
This tool represents a small yet important step forward in democratising accessibility auditing, allowing gadget authors — even those without formal expertise — to identify and rectify errors before they impact our readers.
A Legacy of “Wikiricci” and Community Care
Daimona Eaytoy with the WikiRiccio
The roots of this technical collaboration extend back to 2018 at itWikiCon in Como (Italy), where the “Officina” (the Italian Wikipedia’s technical project) was honoured for its quiet, essential labour, carried out by the smanettoni (hackers) — the tinkerers and wizards who operate behind the scenes to ensure the platform’s gears continue to turn. This community recognition is personified by the Wikiriccio (wiki hedgehog), a physical trophy whose travel history has become something of a legendary saga within the Italian community. Traditionally held in rotation, after years of near-misses, it finally found its way to Daimona Eaytoy during this hackathon, reminding us that accessibility work is also about human connections and shared care.
For us, this light-hearted tradition and award serve as a reminder: behind every accessibility tool or interface fix is a human connection, a shared community-based vision and history, and a commitment to “making the shop run” for the benefit of all users.
Next Steps and Community Involvement
The hackathon session was only the beginning. The outcomes of our session are being synthesised into a formal proposal in the Italian Wikipedia and a Phabricator task to help standardise CSS custom properties and automated linting workflows.
Yet, technology alone cannot solve a cultural challenge. We invite all UI/UX designers, developers, and experienced wiki-editors to join the effort. Whether you are improving the alt text on a high-traffic policy page or helping modernise an old template, your contribution ensures that Wikipedia remains truly accessible, enabling everyone to share in the sum of all knowledge.
A special thanks to the hackathon organisers and all the participants who shared their lived experiences; your insights are what drive these technical improvements forward.
The second cohort of the Africa Wiki Women (AWW) On-Wiki Skills Mentorship Program successfully concluded with a vibrant graduation ceremony celebrating the achievements, growth, and resilience of participants from across Africa. The event marked another major milestone in Africa Wiki Women’s ongoing commitment to empowering women and underrepresented communities through digital literacy, Wikimedia editing skills, mentorship, and leadership development.
About the On-Wiki Skills Mentorship Program
The On-Wiki Skills a 3 month mentorship program designed to equip emerging Wikimedians with practical knowledge and hands-on experience in navigating Wikimedia projects effectively. The program is facilitated through structured mentorship sessions, peer learning sessions , practical exercises, and community engagement activities. Additionally, the mentorship offers a safe learning space for women across the African region to collaborate, learn from each other and strengthen their bond in the wikimedia space creating a rich environment for peer learning, collaboration, and cross-cultural exchange. The second cohort graduated 20 women from 5 African countries, including: Benin, Botswana, Burundi, Cameroon, DR Congo, Ghana, Madagascar, Nigeria, Republic of Congo,Tanzania and Togo.
To ensure effective mentoring, mentees were organized into Anglophone and Francophone cohorts, each supported by mentors who delivered sessions in their respective languages.
Throughout the program, participants received intensive training on key Wikimedia projects including Wikidata, Wikipedia, and Wikimedia Commons. The sessions covered topics such as:
Wikidata principles, policies, item creation, and item improvement
Wikipedia’s Five Pillars, article creation, and notability guidelines
Wikimedia Commons policies, copyright and free licenses, media preparation and uploads, captions, descriptions, metadata, and responsible reuse of Commons content
A Journey of Learning and Transformation
Over the course of the mentorship cycle, participants showed strong commitment throughout the mentorship, consistently attending sessions, completing assignments, and contributing to Wikimedia projects growing from beginners into confident, independent editors.
Baseline and endline report of the second cohort On wiki skill Mentorship program
Mentees contributed in various ways;including:
Creating and improving 316 Wikipedia articles
Uploading 50 media files to Wikimedia Commons
Created 287 Wikidata items
Participating in campaigns and edit-a-thons
Learning effective research and sourcing techniques
Becoming active contributors within Wikimedia communities
These accomplishments reflect the growing impact of mentorship-driven capacity building within the Wikimedia movement, and the Wikimedia Outreach Dashboard created for the mentorship program also provides a detailed record of mentees’ contributions and activities throughout the program.
Highlights from the Graduation Ceremony
The graduation ceremony served as both a celebration and a reflection on the achievements of the cohort. The event featured welcome remarks, mentor appreciations, mentee testimonials, presentations of achievements, and inspiring words from special guest Amanda Jurno.
Speakers commended participants for their resilience, commitment to learning, and willingness to contribute to open knowledge initiatives. Mentors were also recognized for dedicating their time, expertise, and encouragement toward nurturing the next generation of Wikimedians.
Some of the most memorable moments of the ceremony came from mentees sharing personal stories about how the program transformed their confidence, expanded their digital skills, and introduced them to global collaborative communities
Recognizing the Mentors and Organizing Team
The success of the second cohort would not have been possible without the dedication of the guidance of mentors, facilitators, and the organizing team who worked tirelessly behind the scenes to ensure a smooth and impactful learning experience.
Coordinating team: Blessing Ojewuyi Timothy, Andikan Eduok, Adel Bigata, Mammysou17, and the AWW Communications department headed by Pascaline and Adeyinka. Their commitment to mentorship, knowledge sharing helped mentees remain motivated and supported throughout the program.
Looking Ahead
As the second cohort has graduated, the On-Wiki Skills Mentorship Programme continues to empower participants as active contributors and future leaders in the Wikimedia movement. Africa Wiki Women remains committed to creating inclusive spaces where women and marginalized communities can build digital skills and contribute to free knowledge. Congratulations to all graduates for their growth and impact. Follow us on all social media handles at Africa wiki women and stay tuned for the announcement of the next cohort. Be a registered member today and be part of the vibrant community.
From mentee to trainer — how mentorship transformed my Wikimedia journey
From Silence to Curiosity
When I joined the Wikimedia community in 2025, I was excited but unsure of where to begin. The platform felt vast and intimidating, and for a long time I remained passive, watching others contribute while I struggled to find my own entry point.
That uncertainty began to shift when I discovered a community ready to support me. I realized that even small steps — reading articles, observing edits, and asking questions — could open the door to something bigger.
The Turning Point – Africa Wiki Women Mentorship
Everything changed when I was selected as a participant in the On WikiSkills Mentorship Program organized by Africa Wiki Women. Over three months, I received structured, hands‑on training across Wikipedia, Wikidata and Wikimedia Commons.
I learned how to create and edit articles, add structured data, and contribute images. More importantly, I discovered how mentorship can transform hesitation into confidence.
Growth – From Mentee to Trainer
This mentorship gave me more than technical skills — it gave me courage. I moved from being an inactive member to someone who contributes meaningfully to open knowledge.
A major milestone was writing and publishing a Wikipedia article about Samuel Gbadebo Odewumi, a respected Nigerian academic and transport expert. Contributing that article gave me a sense of pride and responsibility, as it ensured that his work and impact are documented for a global audience. It also reminded me that mentorship is not only about learning but about creating knowledge that others can build upon.
The highlight of my journey was becoming a Wikidata trainer, guiding Africa Wiki Women newbies during the April EditHer Africa Contest. In that session, I introduced participants to Wikidata, helping them navigate the same learning curve I once faced. Each edit and training moment became a symbol of empowerment, showing that knowledge grows stronger when shared.
Gratitude and Reflection
I am deeply grateful to the organizers and mentors of Africa Wiki Women for their guidance. Their support helped me find my voice and leadership within the Wikimedia movement.
The testimonial poster created for the program captures this transformation — from mentee to confident contributor — and stands as a reminder of how mentorship can change lives. It is more than an image; it is a symbol of growth, courage, and community.
Looking Ahead
As I continue my journey, I look forward to expanding my contributions and mentoring others. The Program taught me that belonging comes not from knowing everything, but from being willing to learn, share, and grow together.
I now see myself not just as a participant, but as a builder of community — someone who can help others find their own voice in the Wikimedia movement. In particular, I want to encourage more women to come on board, to see themselves as knowledge creators and leaders. Their voices and perspectives are vital, and through initiatives like On WikiSkills Mentorship Program organized by Africa Wiki Women, we can ensure that the Wikimedia projects reflect the richness and diversity of our world.
If you are a woman curious about contributing, now is the time to join us. Your story, your knowledge, and your perspective matter — and together, we can make Wikimedia stronger and more inclusive.
Latest tech news from the Wikimedia technical community. Please tell other users about these changes. Not all changes will affect you. Translations are available.
Updates for editors
The Reader Experience team is conducting an experiment to show the reading lists feature, which is still in development, to logged-out mobile readers to test whether it encourages account creation at a higher rate compared to the watchstar button. The experiment was launched on May 18th on German, Spanish, Italian, Portuguese, Polish, Dutch, Turkish, and Urdu wikis, and it will run for a month.
The Wikimedia Apps team released Phase 1 of the redesigned Home Feed to the Android Beta app. The new Home Feed includes a refreshed “Community” tab and a personalized “For You” tab featuring daily updated reading recommendations. The redesign is part of a broader effort to improve content discovery and create more engaging learning experiences in the Wikipedia apps.
View all 18 community-submitted tasks that were resolved last week. For example, an issue where images could fail to load for some suggested edits on Special:Homepage, leaving the thumbnail stuck in a loading state, has now been fixed. [1]
Wikimedia Korea and Wikimedians of Japan User Group held 「日本・韓国 友好編集月間」”the Korea-Japan Friendship Editing Month” from March 23 to April 17, 2026. This was the second Korea-Japan editathon, following the event held during Asia Month in 2024.
Results report
It appears that 106 articles were created and edited by 29 participants. Thank you all for your participation.
When organizing events like this, I often worry about what will happen if there aren’t enough participants, but Wikipedians are so kind that they always end up joining in before I even realize it.
From the perspective of someone who reviews articles as part of the management team, there are benefits to this kind of opportunity, such as gaining knowledge that you wouldn’t otherwise learn, and satisfying your intellectual curiosity. It’s very educational and good. Wikipedia is a wonderful tool that allows you to share the knowledge you have and the things you want to know with people all over the world.
Article introduction
It would be impossible to introduce all the articles contributed to this editorthon, so I will only introduce the articles that were selected for the April Monthly New/Improved Article Award.
・「老松堂日本行録」…上野ハム…This article was written by Ueno Ham. It is said to be the oldest surviving travelogue of Japan written by a Korean. Apparently, it is something that is studied in high school Japanese history, and when a certain Wikipedian showed me a glossary of Japanese history terms, I was excitedly saying, “This is in there!”
・「朝鮮半島のヒスイ製勾玉」…のりまき…This article was written by Norimaki. I wonder if his experience writing about 「糸魚川のヒスイ」 “Itoigawa Jade” is proving useful. According to the article, it seems that these magatama (comma-shaped beads) may be of Itoigawa origin.
・「朝鮮半島の建築」…犭…This article was written by 犭. It’s surprisingly difficult to summarize such a broad topic into this size, so I think it’s truly impressive.
・「柳川一件」…This is an article I wrote. I will explain more later.
summary
When I first joined Wikipedia (around 2019), my impression was that the only editing event on Wikipedia was “Asia Month,” so I’m very happy to see an increase in these kinds of international events. I’m not very good at socializing, so I’ll leave the initiation of those kinds of conversations to those who are good at it, and I’ll focus on organizing these kinds of events.
Quokka (ESEAP’s mascot character) and souvenirs from Korea(Lin Xiangru, CC BY-SA 4.0, via Wikimedia commons)
Perhaps because of this connection, at the ESEAP Conference in Kaohsiung, Taiwan, I received a gift related to Korea (probably a notepad) from, 韓国のウィキメディアン, a Wikimedian in Korea, and I really like the design of it.
My own view
「朝鮮通信使」狩野安信 “Korean Envoys to Japan” by Kano Yasunobu(I, PHGCOM, CC BY-SA 3.0, via Wikimedia commons)
Now, this is my own view. I thought it would be strange to criticize the merits and demerits of other people’s articles without writing one myself, so I completely revised one of my own articles. It’s an article called「柳川一件」 “The Yanagawa Incident”
I usually write about the various feudal domains of the Edo period, so I didn’t want to stray too far from that if possible. However, 「対馬府中藩」”the Tsushima-Fuchu Domain”, which has some connection to Korea, was too heavy for me. While searching for a subject of just the right difficulty level, I came across this theme.
The Tsushima clan Sou, determined to repair the broken Korea-Japan relations caused by the Bunroku-Keicho War. But couldn’t find a way. They even resorted to tampering official documents to the capital in Japan(apparently they had been doing so regularly before), and managed to send a Korean envoy again, and restore diplomatic relations. However, one of their retainers (Yanagawa), who played a key role in this effort, became dissatisfied with his position and sought independence, ultimately taking the outrageous step of exposing the forgery of the official documents. This family feud is known as the “Yanagawa Incident.” It’s a very interesting story, and I enjoyed writing it. I was thinking, “These guys are tampering with official documents again (lol),” while I was writing it.
It was selected for 月間強化記事賞 “the Monthly Featured Article Award for April”. All’s well that ends well.
There was a time when Wikipedia was the go-to source for information and one of the most trusted tools for research across the world. From students and journalists to researchers and everyday internet users, millions relied on the platform for quick and accessible knowledge. However, as technology continues to evolve, the way people consume information has also changed.
Today, Wikipedia faces growing competition from emerging technologies such as Artificial Intelligence (AI) tools and social media platforms, which now shape how many people search for and engage with information online. As a result, the platform has experienced a decline in page views over the years, raising important questions about its future relevance and visibility in the digital age.
To address these concerns, about 100 Wikimedian affiliates, volunteers, and external experts gathered in Frankfurt am Main from 30 January to 1 February 2026, for the Wikimedia Futures Lab event organised by the Wikimedia movement. The Futures Lab serves as a space for research, experimentation, and forward-thinking conversations on the future of free knowledge.
At a time when technology is rapidly transforming the internet and information-sharing, the event provided an opportunity for participants to reflect on how Wikipedia can continue to remain relevant, visible, and trusted in an increasingly digital and AI-driven world.
Having attended the Wikimedia Futures Lab event, the guests shared their experiences, reflections, and key takeaways from the discussions held in Frankfurt.
“The world around us is changing really fast. When you think about how people trust information online, AI-generated media, new laws, and shifting technologies, it becomes important to understand how these trends affect us as the Wikimedia community,” says Tochi.
Wikipedia vs Digital Age
Despite technological advancement, Wikipedia, once regarded as one of the most trusted digital information platforms, has seen a decline in page views since 2016 as more people turn to AI tools for information. However, it is important to recognise that many AI systems are trained using content from platforms like Wikipedia.
“For example, when you search for something on Google, the AI overview provides a summary alongside references. Very few people actually click on the Wikipedia link for the longer version. This shows that people are still consuming Wikipedia content, but AI tools now act as middlemen,” explains Olubusola.
According to her, this shift means Wikipedia can no longer rely solely on users visiting the platform directly. Instead, it must adapt to changing online habits and find ways to bring information closer to the spaces where audiences already spend their time.
She adds that Wikipedia must adapt by meeting audiences where they already are, bringing information directly to the platforms people use instead of expecting them to always visit the main website.
The solution
The rise of AI and social media has also changed how people consume information. Many users now prefer short-form content over long-form reading because of shrinking attention spans. Since Wikipedia is traditionally a long-form platform, there is growing pressure for it to evolve alongside these changing habits.
For many younger internet users, information is no longer consumed through lengthy articles alone. Videos, creators, podcasts, and short-form explainers are increasingly becoming the preferred way to learn and engage online.
“People are moving away from institution-based information and increasingly relying on personalities. They want direct interaction, and video content makes information easier to consume. As Wikimedia, we need to pay attention to these shifts so we can meet people where they are,” says Ruby.
The Dilemma
Wikimedia exists because of the volunteers who edit and write the content on the platform. While keeping up with technological change is necessary, the movement also faces the challenge of ensuring that technology does not overshadow the human element that has always been at the centre of Wikimedia projects.
As conversations around AI continue to grow, many community members believe the focus should remain on supporting contributors rather than replacing them.
Last year, the Wikimedian community launched its AI Strategy, which clearly showed that AI should not replace the human writers and editors but rather support their work.
After a few years away from Thai Wikipedia, I returned to find that the Main Page had become stagnant. It lacked the dynamic energy a landing page needs. So, my colleagues and I decided to revitalise it—and here is exactly how we did it.
Thai Wikipedia’s Home Page, as of 26 May 2026
Before diving into the details, let me explain the structure of Thai Wikipedia’s Home Page. It was heavily inspired by the original English edition‘s layout, featuring four core content sections:
This Month’s Featured Articles (TMFA): An excerpt of a well-written article (Thai Wikipedia lacks the volume to change this daily like the English site).
Did You Know (DYK): Interesting facts pulled from recently expanded or created articles.
In The News (ITN): Recent global (and occasionally space-related) events.
On This Day (OTD): A look back at historical events on the current date.
When I returned to active editing in mid-2024, I realised these sections were frozen in time. Sometimes, content remained identical for days. After a thorough review, I found the issues were threefold: stagnant content, unpredictable update schedules (except for the strictly automated OTD), and complex, opaque backend procedures for publishing content to the Main Page.
To build a sustainable solution, we had to attack the problem from two angles: community contribution and technical infrastructure.
On the contribution side, we introduced clear, easy-to-follow Standard Operating Procedures (SOPs) to ensure nominators and reviewers wouldn’t feel overwhelmed. We also lifted several legacy constraints that were discouraging newbie and intermediate editors.
On the nerdy side, we introduced a “Nested Transclude Template System” to make pulling content to the main page seamless. No more messy, bespoke coding required. All nominations can now be tracked and recalled without digging through a chaotic page history.
For the less tech-savvy, here is how simple it is now: You no longer need to deal with any messy, complicated coding. As shown in the diagram, everything is built like a set of nesting dolls:
A Diagram to demonstrate a nested template system for Wikipedia. Content like hooks and excerpts are grouped inside date-based templates, which are automatically pulled into the main DYK and TMRA templates.
Write your content: You just write your proposal or excerpt in a standard form.
Name it with the date: You save it inside a specific date format (like YYYY-MM-DD).
The system does the rest: When that day arrives, the Main Page template automatically fetches the correct date’s content and puts it live—completely on its own!
This means no one has to lift a finger to update it manually, and we can track past nominations without digging through a chaotic page history.
Did You Know it’s now easier than ever to nominate your articles?
The first backlog I tackled was the DYK section. There, I crossed paths with Taweethaも, a renowned Thai Wikipedian. That chance encounter inspired a complete revolution of our process. We teamed up to clear backlogs that had been sitting untouched for over six months. Together, we drafted new SOPs and built a backend system to support them—queuing content chronologically by nomination date, enforcing character limits, and scheduling release dates.
Once the system stabilised, we launched a content contest to diversify the topics and test our new workflow under pressure. The campaign was a massive success: 16 contributors created or improved over 90 articles. Crucially, three of those contributors remain highly active “DYK editors” today.
We also noticed that while some nominators were incredibly prolific, they rarely helped review others’ work. To keep the backlog manageable, we implemented a Quid Pro Quo (QPQ) policy, requiring nominators to review a peer’s submission to qualify their own.
Opening the Gates: Allowing Good Articles onto the Main Page
With DYK running smoothly, we turned our attention to TMFA. This section had suffered from a decade-long drought of new Featured Articles (FAs) to showcase. Beyond adapting our new DYK SOPs, we made a major policy shift: we lifted the strict FA constraint and allowed Good Articles (GAs) to be featured. To reflect this, we renamed the section from This Month’s Featured Article to Recommended Articles.
Whilst long-form, high-quality writing requires significantly more energy from contributors—meaning it wasn’t as explosive as the DYK campaign—the initiative still successfully brought 7 brand-new, high-quality articles to the front page from 7 different writers.
A new solution brings a new quirk
An excerpt of Thai Wikipedia’s Home Page on 4 June 2025 showing OTD content from 31 May due to caching issues.
Every new system has its bugs. Just a day into the DYK campaign, a participant noticed that logged-out readers were seeing stale, outdated main page content, while logged-in users saw the updates perfectly.
We spent days hunting for a fix. Thankfully, User:Chlod—a perennial savior of Wikipedia infrastructure—pointed out that the server cache just needed to be manually “purged” (which simply means appending ?action=purge to the URL string).
To automate this, I sat down for some classic “vibe coding” and wrote a Python script. Hosted on Toolforge (Wikimedia’s dedicated server for customised scripts within the Wikimedia Movement) and linked to my bot account, it now runs via a cron job twice a day to keep the page fresh. I also added a secondary feature to the script: it automatically archives the Main Page to the Internet Archive‘s Wayback Machine daily.
For those unfamiliar with the tech jargon, here is the simple version: I asked the AI chatbot, Google Gemini, to help me write a program in the Python language. After testing it repeatedly until I was sure it worked, I uploaded the code to Toolforge—which is essentially a free, 24/7 computer server available to Wikipedia volunteers. I set the server to run my code twice a day to automatically fix the glitch and keep the Main Page fresh. As a bonus, I also programmed it to save a daily copy of the Main Page to the Wayback Machine (a digital archive of the internet) so we always have a historical record.
You might be wondering why I haven’t mentioned ITN or OTD. To be completely honest, I tried to implement similar reforms for OTD, but couldn’t find anyone in the community available to jump in. If you have ideas on how we can spark interest and bring that same magic to the remaining sections, please drop a comment!
Acknowledgements
This transformation wouldn’t have been possible without an incredible support system. Beyond those already mentioned, I want to thank the original architects of the Main Page structure, as well as every single campaign participant who dedicated time to improving Thai Wikipedia. Finally, my deepest respect goes to Taweethaも, whose guidance both on- and off-wiki was invaluable.
For close to two years, my involvement in the Wikimedia ecosystem was mostly technical. I contributed through code during hackathons as a member of Wiki Mentor Africa. I understood the connections among platforms such as Wikipedia, Wikidata, and Wikimedia Commons. I knew their importance, but I also felt there was more I could do. Something was missing in how I was contributing.
I came into the program with one clear goal: to gain a deeper, practical understanding of how to contribute beyond the technical side of Wikimedia. I wanted to move from simply supporting the ecosystem to actively building knowledge within it.
The training opened my eyes to the structure and responsibility behind Wikimedia contributions. I learned that every Wikimedia project is guided by strong principles that protect the quality and reliability of information.
On Wikipedia, content must be notable, verifiable, and supported by reliable sources. On Wikidata, data must be structured, accurate, and referenced. On Wikimedia Commons, files must follow copyright and licensing policies.
These are not just guidelines; they are what make Wikimedia a trusted global knowledge resource.
Learning Through Practice
One of the strongest aspects of the mentorship program was its practical training. The program did not simply explain policies and standards; it required us to apply them through real contributions.
I learned how to properly reference articles, structure content, improve neutrality, and contribute according to Wikimedia standards. At first, this process was challenging. Finding reliable sources, understanding notability requirements, and writing neutrally required patience and attention to detail.
However, through continuous practice and guidance from the trainers, these concepts gradually became clearer and easier to apply.
The trainers also played a major role in making the experience impactful. Complex policies and technical concepts were broken down into simple, understandable steps, making the learning process accessible and encouraging.
Milestones That Changed My Confidence
One major milestone for me during the program was creating two articles and receiving a barnstar in recognition of my contributions.
That moment shifted my confidence completely.
For the first time, I felt that I was no longer just observing how open knowledge is built behind the scenes. I was actively contributing to the preservation and sharing of knowledge myself.
The experience helped me see Wikimedia differently. It became more than a technical ecosystem I contributed to during hackathons. It became a collaborative space where I could directly improve content, document knowledge, and support representation online.
Growing Beyond the Program
Beyond technical editing skills, the mentorship program also changed my perspective on community contribution and leadership.
Looking ahead, I plan to share what I have learned with my community and support the onboarding of new contributors. I am also stepping into a new role as a trainer for an April editathon, which reflects how much this experience has shaped my growth within the Wikimedia movement.
This journey has been both challenging and rewarding. It pushed me to learn, adapt, and contribute more meaningfully.
Wikimedia is more than a platform. It is a collective effort to make knowledge accessible to everyone.
I used to be a reader. Now I’m a contributor. As a Focus Group member, I don’t just consume knowledge, I create it. I am Kewame Veronicah Mompati, a student based in Gaborone, Botswana. I discovered Wikimedia through social media, but I stayed because of purpose. Behind every edit I make is a belief that Botswana’s stories deserve to be seen, cited and preserved.
My first edit
In September 2023, I attended a Wiki edit-a-thon hosted by Wikimedia Community User Group Botswana. I learned how to create an account and translate English Wikipedia articles into Setswana. That first small edit felt huge. Seeing my username appear in the edit history made it real. At that moment, I understood something important: I was no longer just reading Wikipedia; I was part of it.
From occasional edits to consistent contributions
Joining the Focus Group shifted my journey from occasional editing to consistent contribution. Editing stopped being just about correcting or translating text. It became about visibility, representation, and impact. I began contributing across Wikidata and Wikipedia, improving articles, adding reliable sources and translating content into Setswana. So far, I have made 254 contributions on Wikidata, a total of 184 on tn.wikipedia.org, 85 uploads on Wiki Commons, 6 on en.wikipedia and 5 contributions on meta.wikipedia.org helping strengthen information about Botswana in the global knowledge ecosystem. I also expanded into visual storytelling through Wikimedia Commons, uploading photographs from community photo walks. This taught me that knowledge is not only written, it is also visual, cultural, and lived.
Learning beyond editing
Wikimedia didn’t only teach me how to edit. It taught me how knowledge works.I developed stronger digital literacy skills, learned to evaluate and use reliable sources, and began approaching online information more critically. I now understand that citations are not optional, they are essential for credibility and trust. Through co-facilitating training sessions for new editors, I also built confidence in public speaking, teamwork and mentorship. Supporting others, especially young women entering the Wikimedia space, has been one of the most meaningful parts of my journey.
Challenges behind contribution
This journey hasn’t been without challenges. One of the biggest has been discovering how much of Botswana remains undocumented online. I would often try to write about local villages, people or cultural stories, only to find very limited or no reliable sources. Another challenge was balancing editing with academic responsibilities. To stay consistent, I set small but realistic goals: at least three edits per week. I also experienced Wikipedia’s standards first-hand when one of my articles was flagged for deletion. While difficult at the time, it became an important lesson about notability, verifiability, and the importance of strong sourcing.
Why this work matters
Documenting our communities matters because if we do not write ourselves into history, we risk being left out of it. Through this work, I’ve started seeing the world differently. When I visit a village or learn about a local figure, I now think: Does this have a Wikidata item? Is this documented on Wikipedia? Can others learn from this story? I have become more than an editor, I have become a custodian of our stories.
Representation is impact
When someone searches for their hometown and finds nothing, invisibility is reinforced. But when they find well-documented information, images, and history, they find recognition and pride. This is why open knowledge matters.
A call to action
I encourage others to volunteer with Wikimedia. Wikipedia is one of the world’s first points of knowledge discovery, yet African representation remains limited. You do not need advanced technical skills only curiosity and consistency. If you can send a message, you can edit. If you can take a photo, you can upload it to Commons. If you can research, you can contribute sources. My Wikimedia journey is still unfolding. I once thought Wikipedia was written by “them.” Now I know it is written by us. And that changes everything.
The Developer Outreach team is happy to announce that we will be migrating the Tech Blog into Diff. This move will allow us to provide better support and more visibility for the incredible work of the technical community. Diff is the community news and event blog supported by the Movement Communications team. Diff sees about 20,000 visits a month and has 1,200 email subscribers.
What will happen to the Tech Blog content?
All Tech Blog posts will be accessible on Diff, clearly tagged with “techblog”. Old links will automatically redirect to their new location. New posts with a technical focus will be tagged with “techblog” so they will be easily discoverable.You’ll be able to find all techblog posts – old and new – on the landing page at https://diff.wikimedia.org/techblog
When is this happening?
The migration should be complete in April 2026.
How do I submit a blog post with a technical focus?
For now, please hold your post until we complete the migration.
After the migration is done: The process remains the same. For WMF staff, talk to your manager about your interest in writing a blog post so they are not surprised when you ask them to approve it once it is written. For folks outside WMF, if you are part of a team or other larger organization, be sure they are aware and approve. Then, see the Diff submission process and select the category “Technology” and the tag “techblog” when writing your draft. After you submit, the Developer Outreach team will review your draft. When it’s ready to go, we will schedule your post to be published.
We’re excited for the Tech Blog to evolve and thank the Movement Communications team for helping us make this possible!
Today we recognise Thiemo’s broad impact in improving performance of Wikimedia software. From optimizing code across the MediaWiki stack as felt on Wikipedia.org, to speeding up CI for faster developer feedback; this work benefits us every day!
Thiemo Kreuz works in the Technical Wishes team at Wikimedia Deutschland. He did most of this performance work as a paid software developer. “We are free to spend a portion of our time on side projects like these”, Thiemo wrote to us.
Performance as part of a routine
The tools on performance.wikimedia.org are part of building a culture of performance. These tools help you understand how code performs in production and on real devices. These tools empower developers to maintain performance through regular assessment and incremental improvement. Perf matters, because improving performance is an essential step toward equity of access!
We celebrate Thiemo’s tireless efforts with a story about performance as part of a routine, rather than one specific change. We’ll look at a few examples, but there are many other interesting Git commits if you’re curious for more.
Wikitext editor
The CodeMirror extension for MediaWiki provides syntax highlighting, for example, when editing template pages.
“I found a nasty performance issue in CodeMirror’s syntax highlighter for wikitext that was sitting there for a really, really long time”, Thiemo wrote about T270317 and T270237, which would cause your browser to freeze on long articles. “But nobody could figure out why. Answer: Bad regexes with missing boundary assertions.”
VisualEditor template editor
With the WMDE Technical Wishes team, Thiemo worked on VisualEditor’s template dialog and dramatically improved its performance. “This is mostly about lazy-loading parts of the UI”, Thiemo wrote. This matters because the community maintains templates that sometimes define several hundred parameters.
Faster stylesheet compilation
ResourceLoader is the MediaWiki delivery system for frontend styles, scripts, and localisation. It uses the Less.php library for stylesheet compilation. Thiemo heavily optimized the stylesheet parser through native function calls, inlining, and other techniques. This resulted in a 15% reduction in this change, 8% in this change, 5% in another change, and several more changes after that.
The motivation for this work was faster feedback from CI. While we compile only a handful of Less stylesheets during a page view, we have several hundred Less stylesheet files in our codebase. Our CI automatically checks all frontend assets for compilation errors, without needing dedicated unit tests. This speed-up brought us one step closer to realising the 5-minute pipeline.
Codesniffer rules
MediaWiki has extensive static analysis rules that automate and codify things we learned over two decades. Many such rules are implemented using PHP_CodeSniffer and run both locally and in CI via the composer test command. New rules are developed all the time and discussed in Phabricator. These new rules come at a cost.
“I keep coming back to our MediaWiki ruleset for PHPCS to check if it still runs as fast as it used to”, Thiemo wrote. “I find this particularly interesting because it requires a very specific ‘unfair’ type of optimization: We don’t care how slow the unhappy path is when it finds errors, because that’s the exceptional case that typically never happens. But we care a lot about the happy path, because that gets executed over and over again with every CI run.”
Thiemo likes improving low-level libraries and frameworks, such as wikimedia/services and OOUI. “The idea is that even the tiniest of optimizations can make a notable difference, because a piece of library code is executed so often”, Thiemo wrote.
Web Perf Hero award
The Web Perf Hero award is given to individuals who have gone above and beyond to improve the web performance of Wikimedia projects. The initiative started in 2020 and takes the form of a Phabricator badge. You can find past recipients at the Web Perf Hero award page on Wikitech.
How we achieved 20% faster mobile response times, improved SEO, and reduced infrastructure load.
Until now, when you visited a wiki (like en.wikipedia.org), the server responded in one of two ways: a desktop page, or a redirect to the equivalent mobile URL (like en.m.wikipedia.org). This mobile URL in turn served the mobile version of the page from MediaWiki. Our servers have operated this way since 2011, when we deployed MobileFrontend.
Diagram of technical change.
Over the past two months we unified the mobile and desktop domain for all wikis (timeline). This means we no longer redirect mobile users to a separate domain while the page is loading.
We completed the change on Wednesday 8 October after deploying to English Wikipedia. The mobile domains became dormant within 24 hours, which confirms that most mobile traffic arrived on Wikipedia via the standard domains and thus experienced a redirect until now.[1][2]
Why?
Why did we have a separate mobile domain? And, why did we believe that changing this might benefit us?
The year is 2008 and all sorts of websites large and small have a mobile subdomain. The BBC, IMDb, Facebook, and newspapers around the world featured the iconic m-dot domain. For Wikipedia, a separate mobile domain made the mobile experiment low-risk to launch and avoided technical limitations. It became the default in 2011 by way of a redirect.
Fast-forward seventeen years, and much has changed. It is no longer common for websites to have m-dot domains. Wikipedia’s use of it is surprising to our present day audience, and it may decrease the perceived strength of domain branding. The technical limitations we had in 2008 have long been solved, with the Wikimedia CDN having efficient and well-tested support for variable responses under a single URL. And above all, we had reason to believe Google stopped supporting separate mobile domains, which motivated the project to start when it did.
Google used to link from mobile search results directly to our mobile domain, but last year this stopped. This exposed a huge part of our audience to the mobile redirect and regressed mobile response times by 10-20%.[2]
Google supported mobile domains in 2008 by letting you advertise a separate mobile URL. While Google only indexed the desktop site for content, they stored this mobile URL and linked to it when searching from a mobile device.[3] This allowed Google referrals to skip over the redirect.
Google introduced a new crawler in 2016, and gradually re-indexed the Internet with it.[4-7] This new “mobile-first” crawler acts like a mobile device rather than a desktop device, and removes the ability to advertise a separate mobile or desktop link. It’s now one link for everyone! Wikipedia.org was among the last sites Google switched, with May 2024 as the apparent change window.[2] This meant the 60% of incoming pageviews referred by Google, now had to wait for the same redirect that the other 40% of referrals have experienced since 2011.[8]
Persian Wikipedia saw a quarter second cut in the “responseStart” metric from 1.0s to 0.75s.
Unifying our domains eliminated the redirect and led to a 20% improvement in mobile response times.[2] This improvement is both a recovery and a net-improvement because it applies to everyone! It recovers the regression that Google-referred traffic started to experience last year, but also improves response times for all other traffic by the same amount.
The graphs below show how the change was felt worldwide. The “Worldwide p50” corresponds to what you might experience in Germany or Italy, with fast connectivity close to our data centers. The “Worldwide p80” resembles what you might experience in Iran browsing the Persian Wikipedia.
Check Perf report to explore the underlying data and for other regions.
SEO
The first site affected was not Wikipedia but Commons. Wikimedia Commons is the free media repository used by Wikipedia and its sister projects. Tim Starling found in June that only half of the 140 million pages on Commons were known to Google.[9] And of these known pages, 20 million were also delisted due to the mobile redirect. This had been growing by one million delisted pages every month.[10] The cause for delisting turned out to be the mobile redirect. You see, the new Google crawler, just like your browser, also has to follow the mobile redirect.
After following the redirect, the crawler reads our page metadata which points back to the standard domain as the preferred one. This creates a loop that can prevent a page from being updated or listed in Google Search. Delisting is not a matter of ranking, but about whether a page is even in the search index.
Tim and myself disabled the mobile redirect for “Googlebot on Commons” through an emergency intervention on June 23rd. Referrals then began to come back, and kept rising for eleven weeks in a row, until reaching a 100% increase in Google-referrals. From a baseline of 3 million weekly pageviews up to 6 million. Google’s data on clickthroughs shows a similar increase from 1M to 1.8M “clicks”.[9]
Google-referred pageviews in 2025.
Weekly clicks (according to Google Search Console).
We reversed last year’s regression and set a new all-time high. We think there’s three reasons Commons reached new highs:
The redirect consumed half of the crawl budget, thus limiting how many pages could be crawled.[10][11]
Google switched Commons to its new crawler some years before Wikipedia.[12] The index had likely been shrinking for two years already.
Pages on Commons have a sparse link graph. Wikipedia has a rich network of links between articles, whereas pages on Commons represent a photo with an image description that rarely links to other files. This unique page structure makes it hard to discover Commons pages through recursive crawling without a sitemap.
Unifying our domains lifted a ceiling we didn’t know was there!
The MediaWiki software has a built-in sitemap generator, but we disabled this on Wikimedia sites over a decade ago.[13] We decided to enable it for Commons and submitted it to Google on August 6th.[14][15] Google has since indexed 70 million new pages for Commons, up 140% since June.[9]
We also found that less than 0.1% of videos on Commons were recognised by Google as video watch pages (for the Google Search “Videos” tab). I raised this in a partnership meeting with Google Search, and it may’ve been a bug on their end. Commons started showing up in Google Videos a week later.[16][17]
Link sharing UX
When sharing links from a mobile device, such link previously hardcoded the mobile domain. Links shared from a mobile device gave you the mobile site, even when received on desktop. The “Desktop” link in the footer of the mobile site pointed to the standard domain and disabled the standard-to-mobile redirect for you, on the assumption you arrived on the mobile site via the redirect. The “Desktop” link did not remember your choice on the mobile domain itself, and there existed no equivalent mobile-to-standard redirect for when you arrive there. This meant a shared mobile link always presented the mobile site, even after opting-out on desktop.
Everyone now shares the same domain which naturally shows the appropiate version.
There is a long tail of stable referrals from news articles, research papers, blogs, talk pages, and mailing lists that refer to the mobile domain. We plan to support this indefinitely. To limit operational complexity, we now serve these through a simple whole-domain redirect. This has the benefit of retroactively fixing the UX issue because old mobile links now redirect to the standard domain.[18]
This resolves a long-standing bug with workarounds in the form of shared user scripts,[19] browser extensions,[20] and personal scripts.[24]
Infrastructure load
After publishing an edit, MediaWiki instructs the Wikimedia CDN to clear the cache of affected articles (“purge”). It has been a perennial concern from SRE teams at WMF that our CDN purge rates are unsustainable. For every purge from MediaWiki core, the MobileFrontend extension would add a copy for the mobile domain.
Daily purge workload.
After unifying our domains we turned off these duplicate purges, and cut the MediaWiki purge rate by 50%. Over the past weeks the Wikimedia CDN processed approximately 4 billion fewer purges a day. MediaWiki used to send purges at a baseline rate of 40K/second with spikes up to 300K/second, and both have been halved. Factoring in other services, the Wikimedia CDN now receives 20% to 40% fewer purges per second overall, depending on the edit activity.[18]
I don’t have a guestimate for when Google switched Commons to its new crawler. I pinpointed May 2024 as the switch date for Wikipedia based on the new redirect impacting page load times (i.e. a non-zero fetch delay). For Commons, this fetch delay was already non-zero since at least 2018. This suggests Google’s old crawler linked mobile users to Commons canonical domain, unlike Wikipedia which it linked to the mobile domain until last year. Raw perf data: P73601.
Wikipedia is coming up on its 25th birthday, and that would not have been possible without the Wikimedia technical volunteer community. Supporting technical volunteers is crucial to carrying forward Wikimedia’s free knowledge mission for generations to come. In line with this commitment, the Foundation is turning its attention to an important area of developer support—the Wikimedia web (HTTP) APIs.
Both Wikimedia and the Internet have changed a lot over the last 25 years. Patterns that are now ubiquitous standards either didn’t exist or were still in their infancy as the first APIs allowing developers to extend features and automate tasks on Wikimedia projects emerged. In fact, the term “representational state transfer”, better known today as the REST framework, was first coined in 2000, just months before the very first Wikipedia post was published, and only 6 years before the Action API was introduced. Because we preceded what have since become industry standards, our most powerful and comprehensive API solution, the Action API, sticks out as being unlike other APIs – but for good reason, if you understand the history.
Wikimedia APIs are used within Foundation-authored features and by volunteer developers. A common sentiment surfaced through the recent API Listening Tour conducted with a mix of volunteers and Foundation staff is “Wikimedia APIs are great, once you know what you’re doing.” New developers first entering the Wikimedia community face a steep learning curve when trying to onboard due to unfamiliar technologies and complex APIs that may require a deep understanding of the underlying Wikimedia systems and processes. While recognizing the power, flexibility, and mission-critical value that developers created using the existing API solutions, we want to make it easier for developers to make more meaningful contributions faster. We have no plans to deprecate the Action API nor treat it as ‘legacy’. Instead, we hope to make it easier and more approachable for both new and experienced developers to use. We also aim to expand REST coverage to better serve developers who are more comfortable working in those structures.
We are focused on simplifying, modernizing, and standardizing Wikimedia API offerings as part of the Responsible Use of Infrastructure objective in the FY25-26 Annual Plan (see: the WE5.2 key result). Focusing on common infrastructure that encourages responsible use allows us to continue to prioritize reliable, free access to knowledge for the technical volunteer community, as well as the readers and contributors they support. Investing in our APIs and the developer experiences surrounding them will ensure a healthy technical community for years to come. To achieve these objectives, we see three main areas for improving the sustainability of our API offering: simplification, documentation, and communication.
Simplification
To reduce maintenance costs and ensure a seamless developer experience, we are simplifying our API infrastructure and bringing greater consistency across all APIs. Decades of organic growth without centralized API governance led to fragmented, bespoke implementations that now hinder technical agility and standardization. Beyond that, maintaining services is not free; we are paying for duplicative infrastructure costs, some of which are scaling directly with the amount of scraper traffic hitting our services.
In light of the above, we will focus on transitioning at least 70% of our public endpoints to common API infrastructure (see the WE 5.2 key result). Common infrastructure makes it easier to maintain and roll out changes across our APIs, in addition to empowering API authors to move faster. Instead of expecting API authors to build and manage their own solutions for things like routing and rate limiting, we will create centralized tools and processes that make it easier to follow the “golden path” of recommended standards. That will allow centralized governance mechanisms to drive more consistent and sustainable end-user experiences, while enabling flexible, federated API ownership.
An example of simplified internal infrastructure will be introducing a common API Gateway for handling and routing all Wikimedia API requests. Our approach will start as an “invisible gateway” or proxy, with no changes to URL structure or functional behavior for any existing APIs. Centralizing API traffic will make observability across APIs easier, allowing us to make better data-driven decisions. We will use this data to inform endpoint deprecation and versioning, prioritize human and mission-oriented access first, and ultimately provide better support to our developer community.
Centralized management and traffic identification will also allow us to have more consistent and transparent enforcement of our API policies. API policy enforcement enables us to protect our infrastructure and ensure continued access for all. Once API traffic is rerouted through a centralized gateway, we will explore simplifying options for developer identification mechanisms and standardizing how rate limits and other API access controls are applied. The goal is to make it easier for all developers to know exactly what is expected and what limitations apply.
As we update our API usage policies and developer requirements, we will avoid breaking existing community tools as much as possible. We will continue offering low-friction entry points for volunteer developers experimenting with new ideas, lightly exploring data, or learning to build in the Wikimedia ecosystem. But we must balance support for community creativity and innovation with the need to reduce abuse, such as scraping, Denial of Service (DoS) attacks, and other harmful activities. While open, unauthenticated API access for everyone will continue, we will need to make adjustments. To reduce the likelihood and impact of abuse, we may apply stricter rate limits to unauthenticated traffic and more consistent authentication requirements to better match our documented API policy, Robot policy, and API etiquette guidelines, as well as consolidate per-API access guidelines to reduce the likelihood and impact of abuse.
To continue supporting Wikimedia’s technical volunteer community and minimize disruption to existing tools, community developers will have simple ways to identify themselves and receive higher limits or other access privileges. In many cases, this won’t require additional steps. For example, instead of universally requiring new access tokens or authentication methods, we plan to use IP ranges from Wikimedia Cloud Services (WMCS) and User-Agent headers to grant elevated privileges to trusted community tools, approved bots, and research projects.
Documentation
It is essential for any API to enable developers to self-serve their use cases through clear, consistent, and modern documentation experiences. However, Wikimedia API documentation is frequently spread across multiple wiki projects, generated sites, and communication channels, which can make it difficult for developers to find the information they need, when they need it.
To address this, we are working towards a top-requested item coming out of the 2024 developer satisfaction survey: OpenAPI specs and interactive sandboxes for all of our APIs (including conducting experiments to see if we can use OpenAPI to describe the Action API). The MediaWiki Interfaces team began addressing this request through the REST Sandbox, which we released to a limited number of small Wikipedia projects on March 31, 2025. Our implementation approach allows us to generate an OpenAPI specification, which we then use to power a SwaggerUI sandbox. We are also using the OpenAPI specs to automatically validate our endpoints as part of our automated deployment testing, which helps ensure that the generated documentation always matches the actual endpoint behavior.
In addition, the generated OpenAPI spec offers translation support (powered by Translatewiki) for critical and contextual information like endpoint and parameter descriptions. We believe this is a more equitable approach to API documentation for developers who don’t have English as their preferred language. In the coming year, we plan to transition from Swagger UI to a custom Codex implementation for our sandbox experiences, which will enable full translation support for sandbox UI labels and navigation, as well as a more consistent look and feel for Wikimedia developers. We will also expand coverage for OpenAPI specs and sandbox experiences by introducing repeatable patterns for API authors to publish their specs to a single location where developers can easily browse, learn, and make test calls across all Wikimedia API offerings.
Communication
When new endpoints are released or breaking changes are required, we need a better way to keep developers informed. As information is shared through different channels, it can become challenging to keep track of the full picture. Over the next year, we will address this on a few fronts.
First, from a technical change management perspective, we will introduce a centralized API changelog. The changelog will summarize new endpoints, as well as new versions, planned deprecations, and minor changes such as new optional parameters. This will help developers with troubleshooting, as well as help them to more easily understand and monitor the changes happening across the Wikimedia APIs.
In addition to the changelog, we remain committed to consistently communicating changes early and often. As another step towards this commitment, we will provide migration guides and, where needed, provide direct communication channels for developers impacted by the changes to help guarantee a smooth transition. Recognizing that the Wikimedia technical community is split across many smaller communities both on and off-wiki, we will share updates in the largest off-wiki communities, but we will need volunteer support in directing questions and feedback to the right on-wiki pages in various languages. We will also work with communities to make their purpose and audience clearer for new developers so they can more easily get support when they need it and join the discussion with fellow technical contributors.
Over the next few months, we will also launch a new API beta program, where developers are invited to interact with new endpoints and provide feedback before the capabilities are locked into a long-term stable version. Introducing new patterns through a beta program will allow developers to directly shape the future of the Wikimedia APIs to better suit their needs. To demonstrate this pattern, we will start with changes to MediaWiki REST APIs, including introducing API modularization and consistent structures.
What’s Next
We are still in the early stages – we are just making the first steps on the journey to a unified API product offering. But we hope that by this time next year, we will be running towards it together. Your involvement and insights can help us shape a future that better serves the technical volunteers behind our knowledge mission. To keep you informed, we will continue to post updates on mailing lists, Diff, TechBlog, and other technical volunteer communication channels. We also invite you to stay actively engaged: share your thoughts on the WE5 objective in the annual plan, ask questions on the related discussion pages, review slides from the Future of Wikimedia APIs session we conducted at the Wikimedia Hackathon, volunteer for upcoming Listening Tour topics, or come talk to us at upcoming events such as Wikimania Nairobi.
Technical volunteers play an essential role in the growth and evolution of Wikipedia, as well as all other Wikimedia projects. Together, we can make a better experience for developers who can’t remember life before Wikipedia, and make sure that the next generation doesn’t have to live without it. Here’s to another 25 years!
The Campaigns team at WMF has released two features that allow organizers to promote events and WikiProjects on the wikis: Invitation Lists and Collaboration List. These two tools are a part of the CampaignEvents extension, which is available on many wikis.
Invitation Lists
Product overview
Invitation Lists allows organizers to generate a list of people to invite to their WikiProjects, events, or other collaborative activities. It can be accessed by going to Special:GenerateInvitationList, if a wiki has the CampaignEvents extension enabled. You can watch this video demo to see how it works.
It works by looking at a list of articles that an organizer plans to focus on during an activity and then finding users to invite based on the following criteria: the bytes they contributed to the articles, the number of edits they made to the articles, their overall edit count on the wikis, and how recently they have edited the wikis. This makes it easier for organizers to invite people who are already interested in the activity’s topics, hence increasing the likelihood of participation.
With this work, we hope to empower organizers to seek out new audiences. We also hope to highlight the important work done by editors, who may be inspired or touched to receive an invitation to an activity based on their work. However, if someone does not want to receive invitations, they can opt out of being included in Invitation Lists via Preferences.
Technical overview
The “Invitation Lists” feature is part of the CampaignEvents extension for MediaWiki, designed to assist event organizers in identifying and reaching out to potential participants based on their editing activity.
Access and Permissions
Special Pages: The feature introduces two special pages:
Special:GenerateInvitationList: Allows organizers to create new invitation lists.
Special:InvitationList: Displays the generated list of recommended invitees.
User Rights: Access to these pages is restricted to users with the event-organizer right, ensuring that only authorized individuals can generate and view invitation lists.
Invitation List Generation Process
Input Parameters:
List Name: Organizers provide a name for the invitation list.
Target Articles: A list of up to 300 articles relevant to the event’s theme.
The articles will need to be on the wiki of the Invitation List.
Event Page Link: Optionally, a link to the event’s registration page can be included.
Data Collection:
The system analyzes the specified articles to identify contributors.
For each contributor, it gathers metrics such as:
Bytes Added: The total number of bytes the user has added to the articles.
Edit Count: The number of edits made by the user on the specified articles.
Overall Edit Count: The user’s total edit count across the wiki.
Recent Activity: The recency of the user’s edits on the wiki.
Scoring and Ranking:
Contributors are scored based on the collected metrics.
The scoring algorithm assigns weights to each metric to calculate a composite score for each user.
Users are then ranked and categorized into:
Highly Recommended to Invite: Top contributors with high relevance and recent activity.
Recommended to Invite: Contributors with moderate relevance and activity.
Output:
The generated invitation list is displayed on the Special:InvitationList page.
Each listed user includes a link to their contributions page, facilitating further review by the organizer.
Technical Implementation Details
Backend Processing:
The extension utilizes MediaWiki’s job queue system to handle the processing of invitation lists asynchronously, ensuring that the generation process does not impact the performance of the wiki.
Jobs are queued upon submission of the article list and processed in the background.
The articles will need to be on the wiki of the Invitation List, and they can add a maximum of 300 articles.
Data Retrieval:
The extension interfaces with MediaWiki’s revision and user tables to extract the necessary contribution data.
Efficient querying and indexing strategies are employed to handle large datasets and ensure timely processing.
User Preferences and Privacy:
Users have the option to opt out of being included in invitation lists via their preferences.
The extension respects these preferences by excluding opted-out users from the generated lists.
Integration with Event Registration:
If an event page link is provided, the invitation list can be associated with the event’s registration data. This way, we can link their invitation data to their event registration data.
Collaboration List
Product overview
The Collaboration List is a list of events and WikiProjects. It can be accessed by going to SpecialːAllEvents, if a wiki has the CampaignEvents extension enabled.
The Collaboration List has two tabs: “Events” and “Communities.” The Events tab is a global, automated list of all events that use Event Registration. It also has search filters, so you can find events by start and end dates, meeting type (i.e., online, in person, or hybrid), event topic, event wikis, and by keyword searches. You can also find events that are both ongoing (i.e., started before but continue within the selected date range) and upcoming (i.e., events that start within the selected date range).
The Communities tab provides a list of WikiProjects on the local wiki. The WikiProject list is generated by using Wikidata, and it includes: WikiProject name, description, a link to the WikiProject page, and a link to the Wikidata item for the WikiProject. We aim to produce a symbiotic relationship with WikiProjects, in which people can find WikiProjects that interest them, and they can also enhance the Wikidata items for those projects, which in turn improves our project.
Additionally, you can embed the Collaboration List on any wiki page, if the CampaignEvents extension is enabled on that wiki. To do this, you transclude the Collaboration List on a wiki page. You can also choose to customize the Collaboration List through URL parameters, if you want. For example, you can choose to only display a certain number of events or to add formatting. You can read more about this on Help:Extension:CampaignEvents/Collaboration list/Transclusion.
With the Collaboration List, we hope to make it easier for people to find events and WikiProjects that interest them, so more people can find community and make impactful contributions on the wikis together.
Screenshot of the Collaboration List
Technical Overview: Events Tab of Collaboration List
Purpose: Displays a global list of events across all participating wikis.
Data Source: Event data stored centrally in Wikimedia’s X1 database cluster.
Displayed Information:
Event name and description
Event dates (start and end)
Event type (online, in-person, hybrid)
Associated wikis and event topics
Search and Filters:
Date range (start/end)
Meeting type (online, in-person, hybrid)
Event topics and wikis
Keyword search
Ongoing and upcoming event filtering
Technical Implementation:
The CampaignEvents extension retrieves event data directly from centralized tables within the X1 cluster.
Efficient SQL queries and indexing optimize performance for cross-wiki data retrieval.
This implementation ensures quick access and easy discoverability of events from across Wikimedia projects.
Technical Overview: Communities Tab of Collaboration List
Purpose: Displays a list of local WikiProjects on the wiki.
Data Source: Dynamically retrieved from Wikidata via the Wikidata Query Service (WDQS).
Displayed Information:
WikiProject name
Description from Wikidata
Link to the local WikiProject page
Link to the Wikidata item
Performance Optimization:
Query results from WDQS are cached locally using MediaWiki’s caching mechanisms (WANObjectCache).
Cache reduces repeated queries and ensures quick loading times.
Technical Implementation:
The WikimediaCampaignEvents extension retrieves data via SPARQL from WDQS.
The CampaignEvents extension renders the data on Special:AllEvents under the Communities tab.
Extension Communication:
The extensions communicate using MediaWiki’s hook system. The WikimediaCampaignEvents extension provides WikiProject data to the CampaignEvents extension through hook implementations.
This structure enables efficient collaboration between extensions, ensuring clear responsibilities, optimized performance, and simplified discoverability of WikiProjects.
Wikimedia Cloud VPS is a service offered by the Wikimedia Foundation, built using OpenStack and managed by the Wikimedia Cloud Services team. It provides cloud computing resources for projects related to the Wikimedia movement, including virtual machines, databases, storage, Kubernetes, and DNS.
A few weeks ago, in April 2025, we were finally able to introduce IPv6 to the cloud virtual network, enhancing the platform’s scalability, security, and future-readiness. This is a major milestone, many years in the making, and serves as an excellent point to take a moment to reflect on the road that got us here. There were definitely a number of challenges that needed to be addressed before we could get into IPv6. This post covers the journey to this implementation.
The Wikimedia Foundation was an early adopter of the OpenStack technology, and the original OpenStack deployment in the organization dates back to 2011. At that time, IPv6 support was still nascent and had limited implementation across various OpenStack components. In 2012, the Wikimedia cloud users formally requested IPv6 support.
When Cloud VPS was originally deployed, we had set up the network following some of the upstream-recommended patterns:
nova-networks as the engine in charge of the software-defined virtual network
using a flat network topology – all virtual machines would share the same network
using a physical VLAN in the datacenter
using Linux bridges to make this physical datacenter VLAN available to virtual machines
using a single virtual router as the edge network gateway, also executing a global egress NAT – barring some exceptions, using what was called “dmz_cidr” mechanism
In order for us to be able to implement IPv6 in a way that aligned with our architectural goals and operational requirements, pretty much all the elements in this list would need to change. First of all, we needed to migrate from nova-networks into Neutron, a migration effort that started in 2017. Neutron was the more modern component to implement software-defined networks in OpenStack. To facilitate this transition, we made the strategic decision to backport certain functionalities from nova-networks into Neutron, specifically the “dmz_cidr” mechanism and some egress NAT capabilities.
Once in Neutron, we started to think about IPv6. In 2018 there was an initial attempt to decide on the network CIDR allocations that Wikimedia Cloud Services would have. This initiative encountered unforeseen challenges and was subsequently put on hold. We focused on removing the previously backported nova-networks patches from Neutron.
Between 2020 and 2021, we initiated another significant network refresh. We were able to introduce the cloudgw project, as part of a larger effort to rework the Cloud VPS edge network. The new edge routers allowed us to drop all the custom backported patches we had in Neutron from the nova-networks era, unblocking further progress. Worth mentioning that the cloudgw router would use nftables as firewalling and NAT engine.
A pivotal decision in 2022 was to expose the OpenStack APIs to the internet, which crucially enabled infrastructure management via OpenTofu. This was key in the IPv6 rollout as will be explained later. Before this, management was limited to Horizon – the OpenStack graphical interface – or the command-line interface accessible only from internal control servers.
Later, in 2023, following the OpenStack project’s announcement of the deprecation of the neutron-linuxbridge-agent, we began to seriously consider migrating to the neutron-openvswitch-agent. This transition would, in turn, simplify the enablement of “tenant networks” – a feature allowing each OpenStack project to define its own isolated network, rather than all virtual machines sharing a single flat network.
Once we replaced neutron-linuxbridge-agent with neutron-openvswitch-agent, we were ready to migrate virtual machines to VXLAN. Demonstrating perseverance, we decided to execute the VXLAN migration in conjunction with the IPv6 rollout.
We prepared and tested several things, including the rework of the edge routing to be based on BGP/OSPF instead of static routing. In 2024 we were ready for the initial attempt to deploy IPv6, which failed for unknown reasons. There was a full network outage and we immediately reverted the changes. This quick rollback was feasible due to our adoption of OpenTofu: deploying IPv6 had been reduced to a single code change within our repository.
We started an investigation, corrected a few issues, and increased our network functional testing coverage before trying again. One of the problems we discovered was that Neutron would enable the “enable_snat” configuration flag for our main router when adding the new external IPv6 address.
Neutron as the engine in charge of the software-defined virtual network
Ready to use tenant-networks
Using a VXLAN-based overlay network
Using neutron-openvswitch-agent to provide networking to virtual machines
A modern and robust edge network setup
Over time, the WMCS team has skillfully navigated numerous challenges to ensure our service offerings consistently meet high standards of quality and operational efficiency. Often engaging in multi-year planning strategies, we have enabled ourselves to set and achieve significant milestones.
The successful IPv6 deployment stands as further testament to the team’s dedication and hard work over the years. I believe we can confidently say that the 2025 Cloud VPS represents its most advanced and capable iteration to date.
This post is about importing Wikidata into the graph database technology used for hosting the Wikidata Query Service (WDQS). The post includes details on how you can perform your own full Wikidata import to Blazegraph in about a week if you have a nice desktop computer, which was one of the nice takeaways from the analysis.
System utilization around file 20 of 2583 of Wikidata import to local WDQS
Graph databases and Wikidata
Graph databases are a useful technology for data mining relationships between all kinds of things and for enriching knowledge seeking via retrieval-augmented generation (“RAG”) and other AI. Within the Wikimedia content universe, we have a powerful graph database offering called the Wikidata Query Service (“WDQS”) which is based on a mid-2010s technology called Blazegraph.
Wikidata community members model topics you might find on Wikipedia, and this modeling makes it possible to answer all kinds of questions after importing Wikidata’s data into WDQS. Our colleague Trey wrote a nice post describing WDQS that you should check out.
The Wikidata and WDQS architecture spans a number of components and technologies.
Wikidata and Wikidata Query Service high level diagram at it pertains to data flows that may be involved in consumption
Big data growing pains
As Wikidata has grown, the WDQS graph database has become pretty big, with about 16.6 billion records (known as triples) as of this writing, with many intricate relationships between those records that ultimately result in complex and large data structures on disk and in memory. Unfortunately, the WDQS graph database has also become unstable as a result, and this seems to be getting worse as the database gets larger. The last time a data corruption occurred it rippled through the infrastructure and it took about 60 days to reload the graph database to a healthy state across all WDQS graph database servers (part of this had to do with repeated failed imports; hopefully techniques in this here post are instructive to others encountering failed imports).
The long recovery time was a prompt to further enhance the data reload mechanisms and to figure out a way to manage the growth in data volume. Over the course of the last year, the Search Platform Team, which is part of the Data Platform Engineering unit at the Wikimedia Foundation, worked on a project to improve things.
As part of its goal setting, the team determined it should make it possible to support more graph database growth (up to 20 billion rows in total) while being able to recover more reliably and more quickly in the event of a database corruption (within 10 days). The idea being that complex queries are useful – WDQS is one of the most important tools in the Wikidata system – but only if the system is up!
In order to support more database growth, it was pretty clear that either the backend graph database would need to be completely replaced or it would be necessary to split the graph database to buy some time, as the clock had run out on the graph database being stable. A full backend graph database replacement is necessary, but this is a rather complex undertaking and would push timelines out considerably; the replacement is an area for further analysis.
A stopgap solution seemed best. So, the team pursued the approach of splitting the graph database from one monolithic database into separate databases partitioned by two coarse grained knowledge domains: (1) scholarly article entities and (2) everything else. As fate would have it, these two knowledge domains are roughly equivalent in size.
Now, while working through the split of the graph database, although initial testing suggested that it should be possible to achieve a data reload of a graph of 10 billion rows within ten days and reloads for both knowledge domains could run in parallel (thus allowing for 20 billion rows in total), there still wasn’t a lot of room for error. What happens if a graph database corruption happens right when the weekend starts? What if some other sort of server maintenance is blocking the start of a reload for a day or two? We wanted to be certain that we could reload and still have some breathing room to stay within 10 days.
Cumulative Wikidata dump import time using the approach detailed in this post on publicly available dump data
Hardware to the rescue?
From previous investigations it seemed that more powerful servers could speed up data reloads. Obvious, right?
Well, yes and no. It’s a little more complicated. People have tried.
Ghislain Auguste Atemezing analyzed Wikidata imports with Amazon Neptune, which is the commercial SaaS successor to Blazegraph (Blazegraph is no longer actively maintained), as well as other alternatives.
The legacy Blazegraph wiki has some nice guidance on Blazegraph performance optimization, I/O optimization, and query optimization (some of this knowledge is evident in a configuration ticket from as early as 2015 involving one of the original maintainers of Blazegraph). Some of it is still useful and seems to apply, although some changes backported in the JDK plus the sheer scale of Wikidata make some of the settings harder to reason about in practice. Data reloads have become so big and time consuming (think on the order of weeks, not hours) that it is impractical (and expensive) to profile every permutation of hardware configuration and Blazegraph, Java, and operating system configuration.
This said, after noticing that a personal gaming-class machine I bought in 2018 for a machine learning workflow (cross-compiling and applying transfer learning ultimately for an offline Raspberry Pi application) was able to do much faster WDQS imports than what we were seeing on our data center servers, I wanted to understand if there were advances with CPU, memory, and disk in the wild that might point the way to even faster data reloads and wanted to understand better if any software configuration variables could yield bigger performance gains.
Wikidata dump segment import times using the approach detailed in this post
This was explored in T359062, where you’ll find an analysis and running log of import performance on various AWS EC2 configurations, a MacBook Pro (2019 Intel-based), my desktop (2018 Intel-based), our bare metal data center servers, and Amazon Neptune. The takeaways from that analysis were that:
Cloud virtual machines are sufficiently fast for running imports. They may be an option in a pinch.
Removal of CPU governor limits on data center class bare metal servers significantly improved performance. In other words, allowing the CPUs to run at their maximum published clock rates sped up imports.
Removal of CPU governor limits didn’t confer an advantage on prosumer grade computers.
A Blazegraph buffer configuration variable increase significantly improved import speed.
Higher grade hard drives (fast consumer NVMe at home and data center class RAIDed SSDs in production) confer a noticeable performance advantage.
The Amazon Neptune service was by far the fastest option for import. It’s unclear if free or near-free data ingestion observed during the free cloud credit period would extend for additional future imports, though. It is a viable option for imports, but requires additional architectural consideration post an import.
The N-Triples file format (.nt) dramatically improved import speed. It should be (and now, is) used instead of the more complicated Turtle (.ttl) format for imports.
Computing configuration and initial setup
My 2018 personal gaming-class machine with a 6-CPU configuration (up to 4.6 GHz turbo boost) after several years of upgrades has 64 GB of DDR4 RAM and a 4 TB NVMe.
A full Wikidata graph import into Blazegraph took 5.22 days with this configuration and our optimized N-triples files in August 2024.
I had the benefit of pre-split N-triples files produced from our Spark cluster as part of an Airflow DAG that runs weekly, where there are no duplicate lines in the files and there are some additional simplifications compared to the N-triples files produced by legacy jobs in our data dumps infrastructure. If you’re doing this at home without a large Spark cluster, though, you can fetch wikidata-<YYYYMMDD>-all-BETA.nt.bz2 from a datestamped directory in the Wikidata dumps and run some shell commands to prepare files to achieve something similar (do note that the data is less optimized, but it works).
You can at present import somewhat reliably and peformantly with one 4 TB NVMe internal drive and one 2 TB external (or SATA) SSD drive if you are willing to script some file compression to avoid running out of disk. In the example that follows, I assume that you have three drives, though: one 4 TB NVMe drive (let’s say this is your primary drive), one SATA or external 2+ TB SSD (that’s /media/ubuntu/EXTERNAL_DRIVE in the example), and another SATA or external 2+ TB SSD (that’s /media/ubuntu/SOME_OTHER_DRIVE in the example).
The commands
Here are the commands you’ll need to download the Wikidata dump, break it up into smaller files that Blazegraph can handle, and import within a reasonable timeframe.
Note that you’ll need to have a copy of the logback.xml file downloaded to your home directory.
# Download some dependencies sudo apt update sudo apt install bzip2 git openjdk-8-jdk-headless screen git clone https://gerrit.wikimedia.org/r/wikidata/query/rdf cd rdf ./mvnw package -DskipTests sudo mkdir /var/log/wdqs mkdir /home/ubuntu/wtemp # Your username and group may differ from ubuntu:ubuntu sudo chown ubuntu:ubuntu /var/log/wdqs touch /var/log/wdqs/wdqs-blazegraph.log cd dist/target/ tar xzvf service-0.3.*-SNAPSHOT-dist.tar.gz cd service-0.3.*-SNAPSHOT/ cd /media/ubuntu/EXTERNAL_DRIVE mkdir wd cd wd # Run the next multiline command before the weekend - be sure to verify # that your computer will stay awake without reboot. The server throttles # somewhat, so the download takes a while. And the deflate and split-sort-split # also take a while. There are faster ways, but this is easy enough. In case # you were wondering, wget seems to work more reliably than other options. # Torrents do exist for dumps, but be sure to verify their checksums against # dumps.wikimedia.org and verify the date of a given dump. In the following # command pipeline we just print out the checksum for manual verification later, # as it's nice to let this run over a weekend and come back on a Monday to # verify instead of potentially having to wait longer; it usually works fine. date && \ wget https://dumps.wikimedia.org/wikidatawiki/entities/20241216/wikidata-20241216-all-BETA.nt.bz2 && \ date && \ wget https://dumps.wikimedia.org/wikidatawiki/entities/20241216/wikidata-20241216-sha1sums.txt && \ grep wikidata-20241216-all-BETA.nt.bz2 wikidata-20241216-sha1sums.txt && \ sha1sum wikidata-20241216-all-BETA.nt.bz2 && \ date && \ bzcat wikidata-20241216-all-BETA.nt.bz2 | split -d --suffix-length=4 --lines=7812500 --additional-suffix='.nt' - 'wikidata_full_with_duplicates.' && \ date && \ sort wikidata_full_with_duplicates.*.nt --unique --temporary-directory=/home/ubuntu/wtemp | split -d --suffix-length=4 --lines=7812500 --additional-suffix='.ttl.gz' --filter='gzip > $FILE' - 'wikidata_full.' && \ date # Let's head back to where you were: cd ~/rdf/dist/target/service-0.3.*-SNAPSHOT/ mv ~/logback.xml . # Using runBlazegraph.sh like production, change heap from 16g to 31g and # point to logback.xml by updating HEAP_SIZE and LOG_CONFIG to look like so, # without the # comment symbols, of course. # HEAP_SIZE=${HEAP_SIZE:-"31g"} # LOG_CONFIG=${LOG_CONFIG:-"./logback.xml"} vim runBlazegraph.sh # Modify the buffer in RWStore.properties so it looks like this (1M, not 100K), # without the # comment symbol, of course. # com.bigdata.rdf.sail.bufferCapacity=1000000 vim RWStore.properties # Let's get Blazegraph running in the background. screen # Wait a few seconds after running the next command to ensure it's good. ./runBlazegraph.sh # Then CTRL-a-d to leave screen session running in background. # You can chain the following commands together with && \ if you like. # Let's import the first file to make sure it's working (takes about 1 minute). time ./loadData.sh -n wdq -d /media/ubuntu/EXTERNAL_DRIVE/wd -s 0 -e 0 -f 'wikidata_full.%04d.ttl.gz' 2>&1 | tee -a loadData.log # If it worked, let's import another 9 files (maybe another ~10 minutes). time ./loadData.sh -n wdq -d /media/ubuntu/EXTERNAL_DRIVE/wd -s 1 -e 9 -f 'wikidata_full.%04d.ttl.gz' 2>&1 | tee -a loadData.log # Let's see how long it took to import the first ten files, just sum and then # divide by 1000 for seconds (sum / 1000 / 60 / 60 / 24 for days). grep COMMIT loadData.log | cut -f2 -d"=" | cut -f1 -d"m" # Now let's handle the rest of the files. This could take a week or so - again # be sure to verify that your computer will stay awake, without reboot. time ./loadData.sh -n wdq -d /media/ubuntu/EXTERNAL_DRIVE/wd -s 10 -f 'wikidata_full.%04d.ttl.gz' 2>&1 | tee -a loadData.log # Hopefully that worked. Go to http://localhost:9999/bigdata/#query and run the # following query: # SELECT (count(*) as ?ct) WHERE { ?s ?p ?o } # For this example it was 19,827,410,787 with the non-optimized dump. # As of March 2025 you might expect 16.6B for an optimized dump, as here: # https://query.wikidata.org/#select%20%28count%28%2a%29%20as%20%3Fct%29%20where%20%7B%3Fs%20%3Fp%20%3Fo%7D # Celebrate! # Let's close Blazegraph and make a backup of the Blazegraph journal. screen -r # CTRL-c to stop Blazegraph exit # Okay, screen session ended, let's look at the size of the file ls -alh wikidata.jnl cp wikidata.jnl /media/ubuntu/SOME_OTHER_DRIVE/
You’ll notice here I don’t take time to make intermediate backups of the Blazegraph journal file. It’s a good exercise for the reader!
Production, in practice
We were a little surprised that my desktop could perform faster imports than what we were seeing in our data center servers. Our colleague Brian King in Data Platform SRE had a hunch, which turned out to be correct, that we could adjust the CPU governor on the production servers. This helped dramatically on the production servers, and when coupled with the graph split it makes recovery much faster. We don’t need to use the buffer size configuration trick as described above, but we also have that as an option should it become necessary.
Considerations
It would be nice to have no hardware limitations, but there are some practical limitations.
CPU: Although CPU speed increases are still being observed with each new generation of processor, much of the advances in computing have to do with parallelizing computation across more cores. And although WDQS’s graph database holds up relatively well in parallelizing queries across multiple cores, it’s difficult to optimize perfectly for many-cores architecture for data import.
Memory: Although more memory is commonly beneficial to large data operations and intuitively you might expect a graph database to work better with more memory, the manner in which memory is used by running programs can drive performance in surprising ways, ranging from good to bad. WDQS runs on Java technology, and configuration of the Java heap is notoriously challenging for achieving performance without long garbage collection (“GC”) pauses. We deliberately use a 31 GB heap in production for our Blazegraph instances. It’s also important to remember that a large Java heap requires a lot of RAM, which can become expensive.
Nevertheless, more memory can be helpful for filesystem paging operations. Taking the hardware configuration guidance at face value suggests that we would need about 12 TB of memory for the scale of data we have today for an ideal server configuration (we have about 1200 GB of data with about 16.6 billion records). We’re getting by with 128 GB of memory per server, which is much less than 12 TB of memory. We’ve also heard of people using several hundred GB of memory and having reasonable success. More memory would be nice, but today it’s too expensive in a multi-node setup built for redundancy across multiple data centers.
Disk: NVMe disks have brought increased speed to data operations. But backpressure on CPU or memory can also mask what might otherwise be able to manifest with speedier NVMe throughput. NVMEs did show a material performance gain during testing, although presently in production we’re thankfully doing okay with RAIDed data center class SSDs (6 TBs). NVMes would most likely be an improvement in the future in the data center, but they are priced higher for data center quality devices, whereas prosumer grade NVMes for personal computers are reasonably priced; due to the risks of hardware failure we prefer to avoid prosumer grade NVMes in the data center.
Caveats
A few things to remember if you’re running Blazegraph at home:
Be mindful of SERVICE wikibase:mwapi syntax, as it uses external Wikimedia APIs; be sure to avoid rapid repeat queries with this syntax.
Beware of exposing on the network: it doesn’t have the same load balancing and firewall arrangement, as well as other security controls, as the real Wikidata Query Service.
Conclusion
If you are looking to host your own Blazegraph database of Wikidata data without having two graph partitions (i.e., if you want to have the full graph in one partition) you might try the following:
Get a desktop with the fastest CPU possible and acquire a speedy 4 TB NVMe plus 64 GB or more of DDR5 RAM; get a couple larger internal SATA SSD or faster throughput external SSDs if you can, too. As of this writing consumer grade 4 TB NVMes can be had with reasonable price-performance tradeoffs; perhaps 6 GB or 8 GB NVMes with the same level of performance will become available in the next year or two.
Import using the N-Triples format split into multiple files.
Consider scripting the batch import operation to make a backup copy of the graph database for every 100 files imported. That way if your graph database import fails at some point you can troubleshoot and resume from the point of backup. The intermediate backup will slow things down a little but it may save you many days in the end.
If you can’t build or upgrade a desktop of your own, consider use of a cloud server to perform the import, then copy the produced graph database journal file to a more budget friendly computer; remember that in addition to cloud compute and storage costs, there may be data transfer costs.
As you saw up above, there are a few variables in configuration files that you need to update in order to speed an import along.
Production
After splitting the Wikidata graph database in two and removing CPU throttling for Wikidata Query Service production data center nodes, we’re now able to import the WDQS database and catch Blazegraph up to the Wikidata edit stream in less than a week. The way Wikidata updates are applied in the production environment is an interesting topic unto itself, but this diagram gives you an idea of how it works.
Production Wikidata Query Service Kafka Flink-based updater
Acknowledgments
Thank you for reading this post. I’d like to thank the wonderful colleagues in Search Platform, Data Platform SRE, Infrastructure Foundations, Traffic, and Data Center Operations for the solid work on the graph split, and more specifically regarding this post, the support in exploring opportunities to improve performance. I’d like to especially express my gratitude to David Causse (WDQS tech lead and systems thinker), Peter Fischer (Flink-Kafka graph splitter extraordinaire), Erik Bernhardson (thank you for the Airflow environment niceties), Ryan Kemper & Brian King & Stephen Munene & Cathal Mooney (cookbook puppeteers who balance networks and servers with a keyboard), Andrew McAllister (thank you for your analysis of query patterns!), Haroon Shaikh & Renil Thomas (much appreciated on the AWS configs) and Willy Pao & Rob Halsell & Sukhbir Singh (thank you for helping investigate NVMe options). Thank you to my Engineering Director, Olja Dimitrijevic, for encouraging this post, as well as to Tajh Taylor for review. And as ever, major thanks to Guillaume Lederrey, Luca Martinelli, and Lydia Pintscher (WMDE) for partnership in WDQS and Wikidata and its amazing community.
High level overview of Wikidata editing data flows and how it relates to components involved in production
There are numerous industry conferences dedicated to web performance. We have attended and spoken at several of them, and noticed important topics remain underrepresented. While the logistics of organizing a conference is too daunting for our small team, FOSDEM presents an appealing compromise.
The Wikimedia Performance Team organized the inaugural Web Performance devroom at FOSDEM 2020.
FOSDEM is the biggest Free and Open Source software conference in the world. It takes place in Brussels every year, is free to attend, and attracts over 8000 attendees. FOSDEM is known for its many self-organized conference tracks, known as “devrooms”. The logistics are taken care of by FOSDEM, while we focus on programming the content. We ran our own CfP, curate and invite speakers, and emcee the event.
This year saw the completion of two milestones on the MediaWiki Multi-DC roadmap. Multi-DC is a cross-team initiative driven by the Performance Team, to evolve MediaWiki for operation from multiple datacenters. This is motivated by higher resilience, and eliminating steps from switchover procedures. This eases or enables routine maintenance by allowing clusters to be turned off — without a major switchover event.
The Multi-DC initiative has brought about performance and resiliency improvements across the MediaWiki codebases, and at every level of our infrastructure. These gains are effective even in today’s single-DC operation. We resolved long-standing tech debt and improved extension interfaces, which increased developer productivity. We also reduced dependencies, coupling, restructured business logic, and implemented async eventual-consistency solutions.
This year we applied the Multi-DC strategy to MediaWiki’s ChronologyProtector (T254634), and started work on the MainStash DB (T212129).
Today we collect real-user data from pageviews, which alerts when a regression happens, but doesn’t help investigate and fix why. Synthetic testing complements this for desktop browsers, but we have no equivalent for mobile devices. Desktop browsers have an “emulate mobile” option, but DevTools emulation is nothing like real mobile devices.
The goal of the mobile device lab is to find performance regressions on Wikipedia, that are relevant to the experience of our mobile users. Alerts include detailed profiles for investigation, like we do for desktop browsers today.
Starting in 2020, we give out a Web Perf Hero award to individuals who have gone above and beyond to improve site performance. It’s awarded (up to) once a quarter to individuals who demonstrate repeated care and discipline around performance.
Since 2018, we have an on-going survey measuring performance perception on several Wikipedias. You can find the main findings in last year’s blog post. An important take-away was that none of the standard and new metrics we tried, correlate well to real user experience. The “best” metric (page load time) scored a mere 0.14 on the Pearson coefficient scale (from 0 to 1). As such, it remains valuable to survey the real perceived performance, as empirical barometer to validate other performance monitoring.
Data from three cohorts, seen in Grafana. You can see that there’s loose correlation with page load time (“loadEventEnd”). When site performance degrades (time goes up), satisfaction gets worse too (positive percentage goes down). Likewise, when load time improves (yellow goes down), satisfaction improves (green goes up).
“How to Logstash at Wikimedia” (🎥 watch, 📙 slides), explains how we monitor production errors with Logstash dashboards, and demonstrates setting up a triage workflow.
Existing frontend metrics correlated poorly with user-perceived performance. It became clear that the best way to understand perceived performance is still to ask people directly about their experience. We set out to run our own survey to do exactly that, and look for correlations from a range of well-known and novel performance metrics to the lived experience. We partnered with Dario Rossi, Telecom ParisTech, and Wikimedia Research to carry out the study (T187299).
While machine learning failed to explain everything, the survey unearthed many key findings. It gave us newfound appreciation for the old school Page Load Time metric, as the metric that best (or least-terribly) correlated to the real human experience.
The Performance Team has been participating in web standards as individually “invited experts” for a while. We initiated the work for Wikimedia Foundation to become an W3C member organization, and by March 2019 it was official.
As a represented membership organization, we are now collaborating in W3C working groups alongside other major stakeholders to the Web!
In the search for a better user experience metric, we tried out the upcoming Element Timing API for images. This is meant to measure when a given image is displayed on-screen. We enrolled wikipedia.org in the ongoing Google Chrome origin trial for the Element Timing API.
The upcoming Event Timing API is meant to help developers identify slow event handlers on web pages. This is an area of web performance that hasn’t gotten a lot of attention, but its effects can be very frustrating for users.
Via another Chrome origin trial, this experiment gave us an opportunity to gather data, discover bugs in several MediaWiki extensions, and provide early feedback on the W3C Editor’s Draft to the browser vendors designing this API.
We decided to commission the implementation of a browser feature that measures performance from an end-user perspective. The Paint Timing API measures when content appears on-screen for a visitor’s device. This was, until now, a largely Chrome-only feature. Being unable to measure such a basic user experience metric for Safari visitors risks long-term bias, negatively affecting over 20% of our audience. It’s essential that we maintain equitable access and keep Wikimedia sites fast for everyone.
We funded and oversaw implementation of the Paint Timing API in WebKit. We contracted Noam Rosenthal who brings experience in both web standards and upstream WebKit development.
ResourceLoader is Wikipedia’s delivery system for styles, scripts, and localization. It delivers JavaScript code on web pages in two stages. This design prioritizes the user experience through optimal cache performance of HTML and individual modules, and through a consistent experience between page views (i.e. no flip-flopping between pages based on when they were cache). It also achieves a great developer experience by ensuring we don’t mix incompatible versions of modules on the same page, and by ensuring rollout (and rollback) of deployments and complete worldwide in under 10 minutes.
This design rests on the first stage (startup manifest) staying small. We carried out a large-scale audit that shrunk the manifest size back down, and put monitoring and guidelines in place. This work was tracked under T202154.
Identify modules that are unused in practice. This included picking up unfinished or forgotten software deprecations, and removing code for obsolete browser compatibility.
Consolidate modules that did not represent an application entrypoint or logical bundle. Extensions are encouraged to use directories and file splitting for internal organization. Some extensions were registering internal files and directories as public module bundles (like a linker or autoloader), thus growing the startup manifest for all page views.
Shrink the registry holistically through clever math and improved compression.
We wrote new frontend development guides as reference material, enabling developers to understand how each stage of the page load process is impacted by different types of changes. We merged and redirected various older guides in favor of this one.
We published our first AS report, which explores the experience of Wikimedia visitors by their IP network (such as mobile carriers and Internet service providers, also known as Autonomous Systems).
This new monthly report is notable for how it accounts for differences in device type and device performance, because device ownership and content choice is not equally distributed among people and regions. We believe our method creates a fair assessment that focuses specifically on the connectivity of mobile carriers and internet services providers, to Wikimedia datacenters.
The goal is to watch the evolution of these metrics over time, allowing us to identify improvements and potential pain points.
Introduce automatic creation of performance metrics that measure specific chunks of MediaWiki code in core and extensions. Powered by WANObjectCache, via the new WANObjectCache keygroup dashboard in Grafana (T197849).
Develop and launch WikimediaDebug v2 featuring inline performance profiling, dark mode, and Beta Cluster support.
TL;DR:On-wiki search “supports” a lot of “languages”. “Search supports more than 50 language varieties” is a defensible position to take. “Search supports more than 40 languages” is 100% guaranteed! Precise numbers present a philosophical conundrum.
Recently, someone asked the Wikimedia Search Platform Team how many languages we support.
This is a squishy question!
The definition of what qualifies as a language is very squishy. We can try to avoid some of the debate by outsourcing the decision to the language codes we use—different codes equal different languages—though it won’t save us.
Another squishy concept is what we mean by “support”, since the level of language-specific processing provided for each language varies wildly, and even what it means to be “language-specific” is open to interpretation. But before we unrecoverably careen off into the land of philosophy of language, let’s tackle the easier parts of the question.
Full Support
“Full” support for many languages means that we have a stemmer or tokenizer, a stop word list, and we do any necessary language-specific normalization. (See the Anatomy of Search series of blog posts, or the Bare-Bones Basics of Full-Text Search video for technical details on stemmers, tokenizers, stop words, normalization, and more.)
CirrusSearch/Elasticsearch/Lucene
The wiki-specific custom component of on-wiki search is called CirrusSearch, which is built on the Elasticsearch search engine, which in turn is built on the Apache Lucene search library.
Sorani has language code ckb, and it is often called Central Kurdish in English.
Persian and Thai do not have stemmers, but that seems to be because they don’t need them.
Running Count: 33
Elasticsearch 7.10 also has two other language analyzers:
The “Brazilian” analyzer is for Brazilian Portuguese, which is represented by a sub-language code (pt-br). However, the Brazilian analyzer has all separate components, and we do use it for the brwikimedia wiki (“Wiki Movimento Brasil”).
The “CJK” (which stands for “Chinese, Japanese, and Korean”) analyzer only normalizes non-standard half-width and fixed-width characters (ア→ア and A→A), breaks up CJK characters into overlapping bigrams (e.g., ウィキペディア is indexed as ウィ, ィキ, キペ, ペデ, ディ, and ィア), and applies some English stop words. That’s not really “full” support, so we won’t count it here. (We also don’t use it for Chinese or Korean.)
We will count Brazilian Portuguese as a language that we support, but also keep a running sub-tab of “maybe only sort of distinct” language varieties.
We’ll come back to Chinese, Japanese, Korean, and the CJK analyzer a bit later.
Running Count: 33–34 (33 languages + 1 major language variety)
We have found some open source software that does stemming or other processing for particular languages. Some as Elasticsearch plugins, some as stand-alone Java code, and some in other programming languages. We have used, wrapped, or ported as needed to make the algorithms available for our wikis.
We have open-source Serbian, Esperanto, and Slovak stemmers that we ported to Elasticsearch plugins.
There are currently no stop word lists for these languages. However, for a typical significantly inflected alphabetic Indo-European language,† a decent stemmer is the biggest single improvement that can be added to an analysis chain for that language. Stop words are very useful, but general word statistics will discount them even without an explicit stop word list.
Having a stemmer (for a language that needs one) can count as the bare minimum for “full” support.
[†] English is weird in that it is not significantly inflected. Non-Indo-European languages can have very different inflection patterns (like Inuit—so much!—or Chinese—so little!), and non-alphabetic writing systems (like Arabic or Chinese) can have significantly different needs beyond stemming to count as “fully” supported.
For Chinese (Mandarin) we have something beyond the not-so-smart (but much better than nothing!) CJK analyzer provided by Elasticsearch/Lucene. Chinese doesn’t really need a stemmer, but it does need a good tokenizer to break up strings of text without spaces into words. That’s the most important component for Chinese, and we found an open-source plugin to do that. Our particular instantiation of Chinese comes with additional complexity because we allow both Traditional and Simplified characters, often in the same sentence. We have an additional open-source plugin to convert everything to Simplified characters internally.
For Hebrew we found an open-source Elasticsearch plugin that does stemming. It also handles the ambiguity caused by the lack of vowels in Hebrew (by sometimes generating more than one stem).
For Korean, we have another open-source plugin that is much better than the very basic processing provided by the CJK analyzer. It does tokenizing and part-of-speech tagging and filtering.
For Polish and Ukrainian, we found an open-source plugin for each that provides a stemmer and stop word list. They both needed some tweaking to handle odd cases, but overall both were successes.
Running Count: 41–42 (41 languages + 1 major language variety)
Shared Configs
Some languages come in different varieties. As noted before, the distinction between “closely related languages” and “dialects” is partly historical, political, and cultural. Below are some named language varieties with distinct language codes that share language analysis configuration with another language. How you count these is a philosophical question, so we’ll incorporate them into our numerical range.
Egyptian Arabic and Moroccan Arabic use the same configuration as Standard Arabic. Originally they had some extra stop words, but it turned out to be better to use those stop words in Standard Arabic, too. Add two languages/language varieties.
Serbo-Croatian—also called Serbo-Croat, Serbo-Croat-Bosnian (SCB), Bosnian-Croatian-Serbian (BCS), and Bosnian-Croatian-Montenegrin-Serbian (BCMS)—is a pluricentric language with four mutually intelligible standard varieties, namely Serbian, Croatian, Bosnian, and Montenegrin. For various historical and cultural reasons, we have Serbian, Croatian, and Bosnian (but no Montenegrin) wikis, as well as Serbo-Croatian wikis. The Serbian and Serbo-Croatian Wikipedias support Latin and Cyrillic, while the Croatian and Bosnian Wikipedias are generally in Latin script. The Bosnian, Croatian, and Serbo-Croatian wikis use the same language analyzer as the Serbian wikis. Add three languages/language varieties.
Malay is very closely related to Indonesian—close enough that we can use the Elasticsearch Indonesian analyzer for Malay. (Indonesian is a standardized variety of Malay.) Add another language/language variety.
Running Count: 41–48 (41 languages + 7 major language varieties)
Moderate Language-Specific Processing
These languages have some significant language-specific(ish) processing that improves search, while still lacking some obvious component (like a stemmer or tokenizer).
For Japanese, we currently use the CJK analyzer (described above). This is the bare minimum of custom configuration that might be considered “moderate” support. It also stretches the definition of “language-specific”, since bigram tokenizing—which would be useful for many languages without spaces—isn’t really specific to any language, though the decision to apply it is language-specific.
There is a “full” support–level Japanese plugin (Kuromoji) that we tested years ago (and have configured in our code, even), but we decided not to use it because of some problems. We have a long-term plan to re-evaluate Kuromoji (and our ability to customize it for our use cases) and see if we could productively enable it for Japanese.
The Khmer writing system is very complex and—for Historical Technological Reasons™—there are lots of ways to write the same word that all look the same, but are underlyingly distinct sequences of characters. We developed a very complex system that normalizes most sequences to a canonical order. The ICU Tokenizer breaks up Khmer text (which doesn’t use spaces between words) into orthographic syllables, which are very often smaller than words. It’s somewhat similar to breaking up Chinese into individual characters—many larger “natural” units are lost, but all of their more easily detected sub-units are indexed for searching.
This is probably the maximum level of support that counts as “moderate”. It’s tempting to move it to “full” support, but true full support would require tokenizing the Khmer syllables into Khmer words, which requires a dictionary and more complex processing. On the other hand, our support for the wild variety of ways people can (and do!) write Khmer is one place where we currently outshine the big internet search engines.
For Mirandese, we were able to work with a community member to set up elision rules (for word-initial l’, d’, etc., as in some other Romance languages) and translate a Portuguese stop word list.
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3
Azerbaijani, Crimean Tatar, Gagauz, Kazakh, and Tatar have the smallest possible amount of language-specific processing. Like Turkish, they use the uppercase/lowercase pairs İ/i and I/ı, so they have the Turkish version of lowercasing configured.
However, Tatar is generally written in Cyrillic (at least on-wiki). Kazakh is also generally in Cyrillic on-wiki, and the switch to using İ/i and I/ı in the Kazakh Latin script was only made in 2021, so maybe we should count that as half?
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–5
(Un)Intentional Specific Generic Support
Well there’s a noun phrase you don’t see every day—what does it even mean?
Sometimes a language-specific (or wiki community–specific) issue gets generalized to the point where there’s no trace of the motivating source. Conversely, a generic improvement can have an outsized impact on a specific language, wiki, or community.
For example, the Nias language uses lots of apostrophes, and some of the people in its Wikipedia community are apparently more comfortable composing articles in word processors, with the text then being copied to the Nias Wikipedia. Some word processors like to “smarten” quotes and apostrophes, automatically replacing them with the curly variants. This kind of variation makes searching hard. When I last looked (some time ago) it also resulted in Nias Wikipedia having article titles that only differ by apostrophe curliness—I assume people couldn’t find the one so they created the other. Once we got the Phab ticket, we added some Nias-specific apostrophe normalization that fixed a lot of their problems.
Does Nias-specific apostrophe normalization count as supporting Nias? It might arguably fall into the “minimal” category.
About a year later, we cautiously and deliberately tested similar apostrophe normalization for all wikis, and eventually added it as a default, which removed all Nias-specific config in our code.
Does general normalization inspired by a strong need from the Nias Wiki community (but not really inherent to the Nias language) count as supporting Nias? I don’t even know.
At another time, I extended some general normalization upgrades that remove “non-native” diacritics to a bunch of languages, and an unexpectedly large benefit was that it was super helpful in Basque, because Basque searchers often ignore Spanish diacritics on Spanish words, while editors use the correct diacritics in articles, creating a mismatch.
If I hadn’t bothered to do some analysis after going live, I wouldn’t have known about this specific noticeable improvement. On the other hand, if I’d known about the specific problem and there wasn’t a semi-generic solution, I would’ve wanted to implement something Basque-specific to solve it.
Does a general improvement that turns out to strongly benefit Basque count as supporting Basque? I don’t even know! (In practice, this is a slightly philosophical question, since Basque has a stemmer and stopword list, too, so it’s already otherwise on the “full support” list.)
I can’t think of any other language-specific cases that generalized so well—though Nias wasn’t the first or only case of apostrophe-like characters needing to be normalized.
Of course, general changes that were especially helpful to a particular language are easy to miss, if you don’t go looking for them. Even if you do, they can be subtle. The Basque case was much easier for me, personally, to notice, because I don’t speak Basque, but I know a little Spanish, so the Spanish words really stood out as such when looking at the data.
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–5 —I Don’t Even Know: 2+
Vague Categorical Support
It’s easy enough to say that the CJK analyzer supports Japanese (where we are currently using it) and that it would be supporting Chinese and Korean if we were using it for those languages—in small part because it has limited scope, and in large part because it seems specific to Chinese, Japanese, and Korean because of the meaning of “CJK”.
But what about a configuration that is not super specific, but still applied to a subset of languages?
Back in the day, we identified that “spaceless languages” (those whose writing system doesn’t put spaces between words) could benefit from (or be harmed by) specific configurations.
We identified the following languages as “spaceless”. We initially passed on enabling an alternate ranking algorithm (BM25) for them (Phab T152092), but we also deployed the ICU tokenizer for them by default.
Tibetan, Dzongkha, Gan, Japanese, Khmer, Lao, Burmese, Thai, Wu, Chinese, Classical Chinese, Cantonese, Buginese, Min Dong, Cree, Hakka, Javanese, and Min Nan.
14 of those are new.
We eventually did enable BM25 for them, but this list has often gotten special consideration and testing to make sure we don’t unexpectedly do bad things to them when we make changes that seem fine for languages with clearer word boundaries (like Phab T266027).
And what about the case where the “category” we are trying to support is “more or less all of them”? Our recent efforts at cross-wiki “harmonization”—making all language processing that is not language-specific as close to the same as possible on all wikis (see Phab T219550)—was a rising language tide that lifted all/most/many language boats. (An easy to understand example is acronym processing, so that NASA and N.A.S.A. can match more easily. However, some languages—because of their writing systems—have few if any native acronyms. Foreign acronyms (like N.A.S.A.) still show up, though.)
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–5 —I Don’t Even Know: 0–∞
So far we’ve focussed on the most obviously languagey of the language support in Search, which is language analysis. However, there are other parts of our system that support particular wikis in a language-specific way.
Learning to Rank
Learning to Rank (LTR) is a plugin that uses machine learning—based on textual properties and user behavior data—to re-rank search results to move better results higher in the result list.
It makes use of many ranking signals, including making wiki-specific interpretations of textual properties—like word frequency stats, the number of words in a query or document, the distribution of matching terms, etc.
Arguably some of what the model learns is language-specific. Some is probably wiki-specific (say, because Wikipedia titles are organized differently than Wikisource titles), and some may be community-specific (say, searchers search differently on Wikipedia than they do on Wiktionary).
The results are the same or better than our previously hand-tuned ranking, and the models are regularly retrained, allowing them to keep up with changes to the way searchers behave in those languages on those wikis.
Does that count as minimal language-specific support? Maybe? Probably?
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–6 —I Don’t Even Know: 0–∞
Cross-Language Searching
Years ago we worked on a project on some wikis to do language detection on queries that got very few or no results, to see if we could provide results from another wiki. The process was complicated, so we only deployed it to nine of the largest (by search volume) Wikipedias:
Dutch, English, French, German, Italian, Japanese, Portuguese, Spanish, and Russian.
Those are all covered by language analyzers above. However, for each of those wikis, we limited the specific languages that could be identified by the language-ID tool (called TextCat), to maximize accuracy and relevance.
Nine of those are not covered by the language analyzers, and eight are not covered by the LTR plugin: Afrikaans, Breton, Burmese, Georgian, Icelandic, Latin, Tagalog, Telugu, and Urdu. (Vietnamese is covered both by Learning to Rank and TextCat.)
Does sending queries from the largest wikis to other wikis count as some sort of minimal support? Maybe. Arguably. Perhaps.
Running Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–14 —I Don’t Even Know: 0–∞
Conclusions?
What, if any, specific conclusions can we draw? Let’s look again at the list we have so far (even though it is also right above.)
“Final” Count: —Full: 41–48 (41 languages + 7 major language varieties) —Moderate: 3 —Minimal: 4½–14 —I Don’t Even Know: 0–∞
We have good to great support (“moderate” or “full”) for 44 inarguably distinct languages, though it’s very reasonable to claim 51 named language varieties.
The Search Platform team loves to make improvements to on-wiki search that are relevant to all or almost all languages (like acronym handling) or that help all wikis (like very basic parsing for East Asian languages on any wiki). So, how many on-wiki communities does the Search team support? All of them, of course!
Exactly how many languages is that? I don’t even know.
(This blog post is a snapshot from July 2024. If you are from the future, there may be updated details on mediawiki.org.)
Summary: this article shares the experience and learnings of migrating away from Kubernetes PodSecurityPolicy into Kyverno in the Wikimedia Toolforge platform.
Wikimedia Toolforge is a Platform-as-a-Service, built with Kubernetes, and maintained by the Wikimedia Cloud Services team (WMCS). It is completely free and open, and we welcome anyone to use it to build and host tools (bots, webservices, scheduled jobs, etc) in support of Wikimedia projects.
We provide a set of platform-specific services, command line interfaces, and shortcuts to help in the task of setting up webservices, jobs, and stuff like building container images, or using databases. Using these interfaces makes the underlying Kubernetes system pretty much invisible to users. We also allow direct access to the Kubernetes API, and some advanced users do directly interact with it.
Each account has a Kubernetes namespace where they can freely deploy their workloads. We have a number of controls in place to ensure performance, stability, and fairness of the system, including quotas, RBAC permissions, and up until recently PodSecurityPolicies (PSP). At the time of this writing, we had around 3.500 Toolforge tool accounts in the system. We early adopted PSP in 2019 as a way to make sure Pods had the correct runtime configuration. We needed Pods to stay within the safe boundaries of a set of pre-defined parameters. Back when we adopted PSP there was already the option to use 3rd party agents, like OpenPolicyAgent Gatekeeper, but we decided not to invest in them, and went with a native, built-in mechanism instead.
The WMCS team explored different alternatives for this migration, but eventually we decided to go with Kyverno as a replacement for PSP. And so with that decision it began the journey described in this blog post.
First, we needed a source code refactor for one of the key components of our Toolforge Kubernetes: maintain-kubeusers. This custom piece of software that we built in-house, contains the logic to fetch accounts from LDAP and do the necessary instrumentation on Kubernetes to accommodate each one: create namespace, RBAC, quota, a kubeconfig file, etc. With the refactor, we introduced a proper reconciliation loop, in a way that the software would have a notion of what needs to be done for each account, what would be missing, what to delete, upgrade, and so on. This would allow us to easily deploy new resources for each account, or iterate on their definitions.
The initial version of the refactor had a number of problems, though. For one, the new version of maintain-kubeusers was doing more filesystem interaction than the previous version, resulting in a slow reconciliation loop over all the accounts. We used NFS as the underlying storage system for Toolforge, and it could be very slow because of reasons beyond this blog post. This was corrected in the next few days after the initial refactor rollout. A side note with an implementation detail: we stored a configmap on each account namespace with the state of each resource. Storing more state on this configmap was our solution to avoid additional NFS latency.
I initially estimated this refactor would take me a week to complete, but unfortunately it took me around three weeks instead. Previous to the refactor, there were several manual steps and cleanups required to be done when updating the definition of a resource. The process is now automated, more robust, performant, efficient and clean. So in my opinion it was worth it, even if it took more time than expected.
Then, we worked on the Kyverno policies themselves. Because we had a very particular PSP setting, in order to ease the transition, we tried to replicate their semantics on a 1:1 basis as much as possible. This involved things like transparent mutation of Pod resources, then validation. Additionally, we had one different PSP definition for each account, so we decided to create one different Kyverno namespaced policy resource for each account namespace — remember, we had 3.5k accounts.
For developing and testing all this, maintain-kubeusers and the Kyverno bits, we had a project called lima-kilo, which was a local Kubernetes setup replicating production Toolforge. This was used by each engineer in their laptop as a common development environment.
We had planned the migration from PSP to Kyverno policies in stages, like this:
update our internal template generators to make Pod security settings explicit
introduce Kyverno policies in Audit mode
see how the cluster would behave with them, and if we had any offending resources reported by the new policies, and correct them
modify Kyverno policies and set them in Enforce mode
drop PSP
In stage 1, we updated things like the toolforge-jobs-framework and tools-webservice.
In stage 2, when we deployed the 3.5k Kyverno policy resources, our production cluster died almost immediately. Surprise. All the monitoring went red, the Kubernetes apiserver became irresponsibe, and we were unable to perform any administrative actions in the Kubernetes control plane, or even the underlying virtual machines. All Toolforge users were impacted. This was a full scale outage that required the energy of the whole WMCS team to recover from. We temporarily disabled Kyverno until we could learn what had occurred.
This incident happened despite having tested before in lima-kilo and in another pre-production cluster we had, called Toolsbeta. But we had not tested that many policy resources. Clearly, this was something scale-related. After the incident, I went on and created 3.5k Kyverno policy resources on lima-kilo, and indeed I was able to reproduce the outage. We took a number of measures, corrected a few errors in our infrastructure, reached out to the Kyverno upstream developers, asking for advice, and at the end we did the following to accommodate the setup to our needs.:
corrected the external HAproxy kubernetes apiserver health checks, from checking just for open TCP ports, to actually checking the /healthz HTTP endpoint, which more accurately reflected the health of each k8s apiserver.
having a more realistic development environment. In lima-kilo, we created a couple of helper scripts to create/delete 4000 policy resources, each on a different namespace.
greatly over-provisioned memory in the Kubernetes control plane servers. This is, bigger memory in the base virtual machine hosting the control plane. Scaling the memory headroom of the apiserver would prevent it from running out of memory, and therefore crashing the whole system. We went from 8GB RAM per virtual machine to 32GB. In our cluster, a single apiserver pod could eat 7GB of memory on a normal day, so having 8GB on the base virtual machine was clearly not enough. I also sent a patch proposal to Kyverno upstream documentation suggesting they clarify the additional memory pressure on the apiserver.
increased the number of replicas of the Kyverno admission controller to 7, so admission requests could be handled more timely by Kyverno.
I have to admit, I was briefly tempted to drop Kyverno, and even stop pursuing using an external policy agent entirely, and write our own custom admission controller out of concerns over performance of this architecture. However, after applying all the measures listed above, the system became very stable, so we decided to move forward. The second attempt at deploying it all went through just fine. No outage this time 🙂
When we were in stage 4 we detected another bug. We had been following the Kubernetes upstream documentation for setting securityContext to the right values. In particular, we were enforcing the procMount to be set to the default value, which per the docs it was ‘DefaultProcMount’. However, that string is the name of the internal variable in the source code, whereas the actual default value is the string ‘Default’. This caused pods to be rightfully rejected by Kyverno while we figured the problem. We sent a patch upstream to fix this problem.
We finally had everything in place, reached stage 5, and we were able to disable PSP. We unloaded the PSP controller from the kubernetes apiserver, and deleted every individual PSP definition. Everything was very smooth in this last step of the migration.
This whole PSP project, including the maintain-kubeusers refactor, the outage, and all the different migration stages took roughly three months to complete.
For me there are a number of valuable reasons to learn from this project. For one, the scale is something to consider, and test, when evaluating a new architecture or software component. Not doing so can lead to service outages, or unexpectedly poor performances. This is in the first chapter of the SRE handbook, but we got a reminder the hard way 🙂
MediaWiki is the platform that powers Wikipedia and other Wikimedia projects. There is a lot of traffic to these sites. We want to serve our audience in a way that they get the best experience and performance possible. So efficiency of the MediaWiki platform is of great importance to us and our readers.
MediaWiki is a relatively large application with 645,000 lines of PHP code in 4,600 PHP files, and growing! (Reported by cloc.) When you have as much traffic as Wikipedia, working on such a project can create interesting problems.
MediaWiki uses an “autoloader” to find and import classes from PHP files into memory. In PHP, this happens on every single request, as each request gets its own process. In 2017, we introduced support for loading classes from PSR-4 namespace directories (in MediaWiki 1.31). This mechanism involves checking which directory contains a given class definition.
Problem statement
Kunal (@Legoktm) noticed after MediaWiki 1.35, wikis became slower due to spending more time in fstat system calls. Syscalls make a program switch to kernel mode, which is expensive.
We learned that our Autoloader was the one doing the fstat calls, to check file existence. The logic powers the PSR-4 namespace feature, and actually existed before MediaWiki 1.35. But, it only became noticeable after we introduced the HookRunner system, which loaded over 500 new PHP interfaces via the PSR-4 mechanism.
MediaWiki’s Autoloader has a class map array that maps class names to their file paths on disk. PSR-4 classes do not need to be present in this map. Before introducing HookRunner, very few classes in MediaWiki were loaded by PSR-4. The new hook files leveraged PSR-4, exposing many calls to file_exists() for PSR-4 directory searching, in every request. This adds up pretty quickly thereby degrading MediaWiki performance.
See task T274041 on Phabricator for the collaborative investigation between volunteers and staff.
Solution: Optimized class map
Máté Szabó (@TK-999) took a deep dive and profiled a local MediaWiki install with php-excimer and generated a flame graph. He found that about 16.6% of request time was spent in the Autoloader::find() method, which is responsible for finding which file contains a given class.
Figure 1: Flame graph by Máté Szabó.
Checking for file existence during PSR-4 autoloading seems necessary because one namespace can correspond to multiple directories that promise to define some of its classes. The search logic has to check each directory until it finds a class file. Only when the class is not not found anywhere may the program crash with a fatal error.
Máté avoided the directory searching cost by expanding MediaWiki’s Autoloader class map to include all classes, including those registered via PSR-4 namespaces. This solution makes use of a hash-map, where each class maps to one and only one file path on disk, making it a 1-to-1 mapping.
This means, the Autoloader::find() method no longer has to search through the PSR-4 directories. It now knows upfront where each class is, by merely accessing the array from memory. This removes the need for file existence checks. This approach is similar to the autoloader optimization flag in Composer.
Impact
Máté’s optimization significantly reduced response time by optimizing the Autoloader::find() method. This is largely due to the elimination of file system calls.
After deploying the change to MediaWiki appservers in production, we saw a major shift in response times toward faster buckets: a ~20% increase in requests completed within 50ms, and a ~10% increase in requests served under 100ms (T274041#8379204).
Máté analyzed the baseline and classmap cases locally, benchmarking 4800 requests, controlled at exactly 40 requests per second. He found latencies reduced on average by ~12%:
Table 1: Difference in latencies between baseline and classmap autoloader.
Latencies
Baseline
Full classmap
p50 (mean average)
26.2ms
22.7ms (~13.3% faster)
p90
29.2ms
25.7ms (~11.8% faster)
p95
31.1ms
27.3ms (~12.3% faster)
We reproduced Máté’s findings locally as well. On the Git commit right before his patch, Autoloader::find() really stands out.
Figure 2: Profile before optimization.Figure 3: Profile after optimization.
NOTE: We used ApacheBench to load the /wiki/Main_Page URL from a local MediaWiki installation with PHP 8.1 on on Apple M1. We ran it both in a bare metal environment (PHP built-in webserver, 8 workers, no APCU), and in MediaWiki-Docker. We configured our benchmark to run 1000 requests with 7 concurrent requests. The profiles were captured using Excimer with a 1ms interval. The flame graphs were generated with Speedscope, and the box plots were created with Gnuplot.
In Figure 4 and 5, the “After” box plot has a lower median than the “Before” box plot. This means there is a reduction in latency. Also, the standard deviation in the “After” scenario shrunk, which indicates that responses were more consistently fast (not only on average). This increases the percentage of our users that have an experience very close to the average response time of web requests. Fewer users now experience an extreme case of web response slowness.
Figure 4: Boxplot for requests on bare metal.Figure 5: Boxplot for requests on Docker.
Web Perf Hero award
The Web Perf Hero award is given to individuals who have gone above and beyond to improve the web performance of Wikimedia projects. The initiative is led by the Performance Team and started mid-2020. It is awarded quarterly and takes the form of a Phabricator badge.
You might have already heard the buzz: the Wikimedia Hackathon is gearing up for an incredible event in Tallinn, Estonia, from May 3rd to 5th, 2024. Now, we’re thrilled to announce that the Registration form, which also includes an optional Scholarship application, is officially open until FridayJanuary 5th 2024.
Participation in the in-person Wikimedia Hackathon in Tallinn is contingent upon registration. The registration portal will remain accessible until we hit our venue’s capacity, which is approximately set at 220 participants. Here’s the exciting part: the event itself is entirely free of charge ensuring that everyone has an opportunity to join us for this fantastic experience. Please note that participants are required to make individual travel arrangements unless they have been awarded a scholarship. For comprehensive details about the scholarship process, committee, and eligibility criteria, please visit the dedicated page.
The registration and scholarship application form is powered by Pretix, an open-source third-party service, which may introduce additional terms. If you have inquiries regarding privacy and data handling, consult the privacy statement for more information.
Seize the Opportunity: Apply Now!
Register, apply for a scholarship, and join the vibrant technical community dedicated to making a difference and shaping the future of Wikimedia’s Technical Ecosystem.
Stay Connected: Join the Conversation
As the excitement builds, stay connected with the Wikimedia community. Engage in discussions, share your ideas, and connect with fellow participants on the talk page and explore the various channels mentioned here. Follow the event updates, announcements, and get ready for an enriching experience that goes beyond coding — it’s about building connections and leaving a lasting impact on the Wikimedia Technical projects.
Should you have any questions or encounter issues related to the registration form or scholarship application, don’t hesitate to reach out to the organizing team. You can connect with us via the talk page or through email at hackathon@wikimedia.org. We’re here to support you every step of the way!
We are thrilled to share the exciting news that the 2024 Wikimedia Hackathon is scheduled to unfold in the captivating city of Tallinn, Estonia, from May 3rd – 5th 2024!
A Celebration of Innovation and Collaboration
The Wikimedia Hackathon is not just an annual hacking event, it’s a celebration of innovation and collaboration, uniting the global Wikimedia technical community in a dynamic gathering focused on connection, innovation, and exploration. At this event, technical contributors hailing from all corners of the globe converge with a shared mission: to enhance the technological infrastructure and software that underpins and empowers Wikimedia projects.
The theme for this edition aligns with last year’s, emphasizing the gathering of individuals who have a track record of contributing to the technical aspects of Wikimedia projects. We’re looking for those who are well-versed in navigating the technical ecosystem and are adept at working autonomously or collaborating effectively on projects.
How to Get Involved
Participating in the Hackathon is easy! Simply mark your calendar for May 3rd – 5th, 2024, and register to attend. Stay tuned for registration and scholarship details, which will be announced on Monday November 27th 2023 on our MediaWiki page and social media channels.
Spread the Word!
Help us make the Hackathon a massive success by spreading the word. Share this announcement with community members,, and anyone who shares your passion for Wikimedia Technical Projects. Let’s make this gathering of brilliant minds an unforgettable experience!
The new “Excimer UI” option in WikimediaDebug generates flame graphs. What are flame graphs, and when do you need this?
A flame graph visualizes a tree of function calls across the codebase, and emphasizes the time each function spends. In 2014, we introduced Arc Lamp to help detect and diagnose performance issues in production. Arc Lamp samples live traffic and publishes daily flame graphs. This same diagnostic power is now available on-demand to debug sessions!
Debugging until now
WikimediaDebug is a browser extension for Firefox and Chromium-based browsers. It helps stage deployments and diagnose problems in backend requests. It can pin your browser to a given data center and server, send verbose messages to Logstash, and… capture performance profiles!
Our main debug profiler has been XHGui. XHGui is an upstream project that we first deployed in 2016. It’s powered by php-tideways under the hood, which favors accuracy in memory and call counts. This comes at the high cost of producing wildly inaccurate time measurements. The Tideways data model also can’t represent a call tree, needed to visualize a timeline (learn more, upstream change). These limitations have led to misinterpretations and inconclusive investigations. Some developers work around this manually with time-consuming instrumentation from a production shell. Others might repeatedly try fixing a problem until a difference is noticeable.
Screenshot of XHGui.
Accessible performance profiling
Our goal is to lower the barrier to performance profiling, such that it is accessible to any interested party, and quick enough to do often. This includes reducing knowledge barriers (internals of something besides your code), and mental barriers (context switch).
You might wonder (in code review, in chat, or reading a mailing list) why one thing is slower than another, what the bottlenecks are in an operation, or whether some complexity is “worth” it?
With WikimediaDebug, you flip a switch, find out, and continue your thought! It is part of a culture in which we can make things faster by default, and allows for a long tail of small improvements that add up.
Example: In reviewing a change, which proposes adding caching somewhere, I was curious. Why is that function slow? I opened the feature and enabled WikimediaDebug. That brought me to an Excimer profile where you can search (ctrl-F) for the changed function (“doDomain”). We find exactly how much time is spent in that particular function. You can verify our results, or capture your own!
Flame graph in Excimer UI via Speedscope (by Jamie Wong, MIT License).
What: Production vs Debugging
We measure backend performance in two categories: production and debugging.
“Production” refers to live traffic from the world at large. We collect statistics from MediaWiki servers, like latency, CPU/memory, and errors. These stats are part of the observability strategy and measure service availability (“SLO”). To understand the relationship between availability and performance, let’s look at an example. Given a browser that timed out after 30 seconds, can you tell the difference between a response that will never arrive (it’s lost), and a response that could arrive if you keep waiting? From the outside, you can’t!
When setting expectations, you thus actually define both “what” and “when”. This makes performance and availability closely intertwined concepts. When a response is slower than expected, it counts toward the SLO error budget. We do deliver most “too slow” responses to their respective browser (better than a hard error!). But above a threshold, a safeguard stops the request mid-way, and responds with a timeout error instead. This protects us against misuse that would drain web server and database capacity for other clients.
These high-level service metrics can detect regressions after software deployments. To diagnose a server overload or other regression, developers analyze backend traffic to identify the affected route (pageview, editing, login, etc.). Then, developers can dig one level deeper to function-level profiling, to find which component is at fault. On popular routes (like pageviews), Arc Lamp can find the culprit. Arc Lamp publishes daily flame graphs with samples from MediaWiki production servers.
Production profiling is passive. It happens continuously in the background and represents the shared experience of the public. It answers: What routes are most popular? Where is server time generally spent, across all routes?
“Debug” profiling is active. It happens on-demand and focuses on an individual request—usually your own. You can analyze any route, even less popular ones, by reproducing the slow request. Or, after drafting a potential fix, you can use debugging tools to stage and verify your change before deploying it worldwide.
These “unpopular” routes are more common than you might think. Wikipedia is among the largest sites with ~8 million requests per minute. About half a million are pageviews. Yet, looking at our essential workflows, anything that isn’t a pageview has too few samples for real-time monitoring. Each minute we receive a few hundred edits. Other workflows are another order of magnitude below that. We can take all edits, reviews of edits (“patrolling”), discussion replies, account blocks, page protections, etc; and their combined rate would be within the error budget of one high-traffic service.
Excimer to the rescue
Tim Starling on our team realized that we could leverage Excimer as the engine for a debug profiler. Excimer is the production-grade PHP sampling profiler used by Arc Lamp today, and was specifically designed for flame graphs and timelines. Its data model represents the full callstack.
Remember that we use XHGui with Tideways, which favors accurate call counts by intercepting every function call in the PHP engine. That costly choice skews time. Excimer instead favors low-overhead, through a sampling interval on a separate thread. This creates more representative time measures. Re-using Excimer felt obvious in retrospect, but when we first deployed the debug services in 2016, Excimer did not yet exist. As a proof of concept, we first created an Excimer recipe for local development.
How it works
After completing the proof of concept, we identified four requirements to make Excimer accessible on-demand:
Capture the profiling information,
Store the information,
Visualize the profile in a way you can easily share or link to,
Discover and control it from an interface.
We took the capturing logic as-is from the proof of concept, and bundled it in mediawiki-config. This builds on the WikimediaDebug component, with an added conditional for the “excimer” option.
To visualize the data we selected Speedscope, an interactive profile data visualization tool that creates flame graphs. We did consider Brendan Gregg’s original flamegraph.pl script, which we use in Arc Lamp. flamegraph.pl specializes in aggregate data, using percentages and sample counts. This is great for Arc Lamp’s daily summaries, but when debugging a single request we actually know how much time has passed. It would be more intuitive to developers if we presented the time measurements, instead of losing that information. Speedscope can display time.
We store each captured profile in a MySQL key-value table, hosted in the Foundation’s misc database cluster. The cluster is maintained by SRE Data Persistence, and also hosts the databases of Gerrit, Phabricator, Etherpad, and XHGui.
Freely licensed software
We use Speedscope as the flame graph visualizer. Speedscope is an open source project by Jamie Wong. As part of this project we upstreamed two improvements, including a change to bundle a font rather than calling on a third-party CDN. This aligns with our commitment to privacy and independence.
The underlying profile data is captured by Excimer, a low-overhead sampling profiler for PHP. We developed Excimer in 2018 for Arc Lamp. To make the most of Speedscope’s feature set, we added support for time units and added the Speedscope JSON format as a built-in output type for Excimer.
We added Excimer to the php.net registry and submitted it to major Linux package managers (Debian, Ubuntu, Sury, and Remi’s RPM). Special thanks to Kunal Mehta as Debian Developer and fellow Wikimedian who packaged Excimer for Debian Linux. These packages make Excimer accessible to MediaWiki contributors and their local development environment (e.g. MediaWiki-Docker).
Our presence in the Debian repository carries special meaning. Presence in the Debian repository signals trust, stability, and confidence in our software to the free software ecosystem. For example, we were pleased to learn that Sentry adopted Excimer to power their Sentry Profiling for PHP service!
Try it!
If you haven’t already, install WikimediaDebug in your Firefox or Chrome browser.
Navigate to any article on Wikipedia.
Set the widget to On, with the “Excimer UI” checked.
Reload the page.
Click the “Open profile” link in the WikimediaDebug popup.
Accessible debugging tools empower you to act on your intuitions and curiosities, as part of a culture where you feel encouraged to do so. What we want to avoid is filtering these intuitions down to big incidents only, where you can justify hours of work, or depend on specialists.
Learn why we transitioned the MediaWiki platform to serve traffic from multiple data centers, and the challenges we faced along the way.
Wikimedia Foundation provides access to information for people around the globe. When you visit Wikipedia, your browser sends a web request to our servers and receives a response. Our servers are located in multiple geographically separate datacenters. This gives us the ability to quickly respond to you from the closest possible location.
You can find out which data center is handling your requests by using the Network tab in your browser’s developer tools (e.g. right-click -> Inspect element -> Network). Refresh the page and click the top row in the table. In the “x-cache” response header, the first digit corresponds to a data center in the above map.
In the example above, we can tell from the 4 in “cp4043”, that San Francisco was chosen as my nearest caching data center. The cache did not contain a suitable response, so the 2 in “mw2393” indicates that Dallas was chosen as the application data center. These are the ones where we run the MediaWiki platform on hundreds of bare metal Apache servers. The backend response from there is then proxied via San Francisco back to me.
Why multiple data centers?
Our in-house Content Delivery Network (CDN) is deployed in multiple geographic locations. This lowers response time by reducing the distance that data must travel, through (inter)national cables and other networking infrastructure from your ISP and Internet backbones. Each caching data center that makes up our CDN, contains cache servers that remember previous responses to speed up delivery. Requests that have no matching cache entry yet, must be forwarded to a backend server in the application data center.
If these backend servers are also deployed in multiple geographies, we lower the latency for requests that are missing from the cache, or that are uncachable. Operating multiple application data centers also reduces organizational risk from catastrophic damage or connectivity loss to a single data center. To achieve this redundancy, each application data center must contain all hardware, databases, and services required to handle the full worldwide volume of our backend traffic.
Multi-region evolution of our CDN
Wikimedia started running its first datacenter in 2004, in St Petersburg, Florida. This contained all our web servers, databases, and cache servers. We designed MediaWiki, the web application that powers Wikipedia, to support cache proxies that can handle our scale of Internet traffic. This involves including Cache-Control headers, sending HTTP PURGE requests when pages are edited, and intentional limitations to ensure content renders the same for different people. We originally deployed Squid as the cache proxy software, and later replaced it with Varnish and Apache Traffic Server.
In 2005, with only minimal code changes, we deployed cache proxies in Amsterdam, Seoul, and Paris. More recently, we’ve added caching clusters in San Francisco, Singapore, and Marseille. Each significantly reduces latency from Europe and Asia.
Adding cache servers increased the overhead of cache invalidation, as the backend would send an explicit PURGE request to each cache server. After ten years of growth both in Wikipedia’s edit rate and the number of servers, we adopted a more scalable solution in 2013 in the form of a one-to-many broadcast. This eventually reaches all caching servers, through a single asynchronous message (based on UDP multicast). This was later replaced with a Kafka-based system in 2020.
When articles are temporarily restricted, “View source” replaces the familiar “Edit” link for most readers.
The traffic we receive from logged-in users is only a fraction of that of logged-out users, while also being difficult to cache. We forward such requests uncached to the backend application servers. When you browse Wikipedia on your device, the page can vary based on your name, interface preferences, and account permissions. Notice the elements highlighted in the example above. This kind of variation gets in the way of whole-page HTTP caching by URL.
Our highest-traffic endpoints are designed to be cacheable even for logged-in users. This includes our CSS/JavaScript delivery system (ResourceLoader), and our image thumbnails. The performance of these endpoints is essential to the critical path of page views.
Multi-region for application servers
Wikimedia Foundation began operating a secondary data center in 2014, as contingency to facilitate a quick and full recovery within minutes in the event of a disaster. We excercise full switchovers annually, and we use it throughout the year to ease maintenance through partial switchover of individual backend services.
Actively serving traffic from both data centers would add advantages over a cold-standby system:
Requests are forwarded to closer servers, which reduces latency.
Traffic load is spread across more hardware, instead of half sitting idle.
No need to “warm up” caches in a standby data center prior to switching traffic from one data center to another.
With multiple data centers in active use, there is institutional incentive to make sure each one can correctly serve live traffic. This avoids creation of services that are configured once, but not reproducible elsewhere.
We drafted several ideas into a proposal in 2015, to support multiple application data centers. Many components of the MediaWiki platform assumed operating from one backend data center. Such as assuming that a primary database is always reachable for querying, or that deleting a key from “the” Memcache cluster suffices to invalidate a cache. We needed to adopt new paradigms and patterns, deploy new infrastructure, and update existing components to accommodate these. Our seven-year journey ended in 2022, when we finally enabled concurrent use of multiple data centers!
The biggest changes that made this transition possible are outlined below.
HTTP verb traffic routing
MediaWiki was designed from the ground up to make liberal use of relational databases (e.g. MySQL). During most HTTP requests, the backend application makes several dozen round trips to its databases. This is acceptable when those databases are physically close to the web servers (<0.2ms ping time). But, this would accumulate significant delays if they are in different regions (e.g. 35ms ping time).
MediaWiki is also designed to strictly separate primary (writable) from replica (read-only) databases. This is essential at our scale. We have a CDN and hundreds of web servers behind it. As traffic grows, we can add more web servers and replica database servers as-needed. But, this requires that page views don’t put load on the primary database server — of which there can be only one! Therefore we optimize page views to rely only on queries to replica databases. This generally respects the “method” section of RFC 9110, which states that requests that modify information (such as edits) use HTTP POST requests, whereas read actions (like page views) only involve HTTP GET (or HTTP HEAD) requests.
The above pattern gave rise to the key idea that there could be a “primary” application datacenter for “write” requests, and “secondary” data centers for “read” requests. The primary databases reside in the primary datacenter, while we have MySQL replicas in both data centers. When the CDN has to forward a request to an application server, it chooses the primary datacenter for “write” requests (HTTP POST) and the closest datacenter for “read” requests (e.g. HTTP GET).
We cleaned up and migrated components of MediaWiki to fit this pattern. For pragmatic reasons, we did make a short list of exceptions. We allow certain GET requests to always route to the primary data center. The exceptions require HTTP GET for technical reasons, and change data at the same low frequency as POST requests. The final routing logic is implemented in Lua on our Apache Traffic Server proxies.
Media storage
Our first file storage and thumbnailing infrastructure relied on NFS. NetApp hardware provided mirroring to standby data centers.
By 2012, this required increasingly expensive hardware and proved difficult to maintain. We migrated media storage to Swift, a distributed file store.
As MediaWiki assumed direct file access, Aaron Schulz and Tim Starling introduced the FileBackend interface to abstract this. Each application data center has its own Swift cluster. MediaWiki tries writes to both clusters, and the “swiftrepl” background service manages consistency. When our CDN finds thumbnails absent from its cache, it forwards requests to the nearest Swift cluster.
Job queue
MediaWiki features a job queue system since 2009, for performing background tasks. We took our Redis-based job queue service, and migrated to Kafka in 2017. With Kafka, we support bidirectional and asynchronous replication. This allows MediaWiki to quickly and safely queue jobs locally within the secondary data center. Jobs are then relayed to and executed in the primary data center, near the primary databases.
The bidirectional queue helps support legacy features that discover data updates during a pageview or other HTTP GET request. Changing each of these features was not feasible in a reasonable time span. Instead, we designed the system to ensure queueing operations are equally fast and local to each data center.
In-memory object cache
MediaWiki uses Memcached as an LRU key-value store to cache frequently accessed objects. Though not as efficient as whole-page HTTP caching, this very granular cache is suitable for dynamic content.
Some MediaWiki extensions assumed that Memcached had strong consistency guarantees, or that a cache could be invalidated by setting new values at relevant keys when the underlying data changes. Although these assumptions were never valid, they worked well enough in a single data center.
We introduced WANObjectCache as a simple yet robust interface in MediaWiki. It takes care of dealing with multiple independent data centers. The system is backed by mcrouter, a Memcached proxy written by Facebook. WANObjectCache provides two basic functions: getWithSet and delete. It uses cache-aside in the local data center, and broadcasts invalidation to all data centers. We’ve migrated virtually all Memcached interactions in MediaWiki to WANObjectCache.
Parser cache
Most of a Wikipedia page is the HTML rendering of the editable content. This HTML is the result of parsing wikitext markup and expanding template macros. MediaWiki stores this in the ParserCache to improve scalability and performance. Originally, Wikipedia used its main Memcached cluster for this. In 2011, we added MySQL as the lower tier key-value store. This improved resiliency from power outages and simplified Memcached maintenance. ParserCache databases use circular replication between data centers.
Ephemeral object stash
The MainStash interface provides MediaWiki extensions on the platform with a general key-value store. Unlike Memcached, this is is a persistent store (disk-backed, to survive restarts) and replicates its values between data centers. Until now, in our single data center setup, we used Redis as our MainStash backend.
In 2022 we moved this data to MySQL, and replicate it between data centers using circular replication. Our access layer (SqlBagOStuff) adheres to a Last-Write-Wins consistency model.
Login sessions were similarly migrated away from Redis, to a new session store based on Cassandra. It has native support for multi-region clustering and tunable consistency models.
Reaping the rewards
Most multi-DC work took the form of incremental improvements and infrastructure cleanup, spread over several years. While we did find latency redunction on some of the individual changes, we mainly looked out for improvements in availability and reliability.
The final switch to “turn on” concurrent traffic to both application data centers was the HTTP verb routing. We deployed it in two stages. The first stage applied the routing logic to 2% of web traffic, to reduce risk. After monitoring and functional testing, we moved to the second stage: route 100% of traffic.
We reduced latency of “read” requests by ~15ms for users west of our data center in Carrollton (Texas, USA). For example, logged-in users within East Asia. Previously, we forwarded their CDN cache-misses to our primary data center in Ashburn (Virginia, USA). Now, we could respond from our closer, secondary, datacenter in Carrollton. This improvement is visible in the 75th percentile TTFB (Time to First Byte) graph below. The time is in seconds. Note the dip after 03:36 UTC, when we deployed the HTTP verb routing logic.
For over 15 years, the Wikimedia Foundation has provided public dumps of the content of all wikis. They are not only useful for archiving or offline reader projects, but can also power tools for semi-automated (or bot) editing such as AutoWikiBrowser. For example, these tools comb through the dumps to generate lists of potential spelling mistakes in articles for editors to fix. For researchers, the dumps have become an indispensable data resource (footnote: Google Scholar lists more than 16,000 papers mentioning the word “Wikipedia dumps”). Especially in the area of natural language processing, the use of Wikipedia dumps has become almost ubiquitous with the advancement of large language models such as GPT-3 (and thus by extension also the recently published ChatGPT) or BERT. Virtually all language models are trained on Wikipedia content, especially multilingual models which rely heavily on Wikipedia for many lower-resourced languages.
Over time, the research community has developed many tools to help folks who want to use the dumps. For instance, the mwxml Python library helps researchers work with the large XML files and iterate through the articles within them. Before analyzing the content of the individual articles, researchers must usually further preprocess them, since they come in wikitext format. Wikitext is the markup language used to format the content of a Wikipedia article in order to, for example, highlight text in bold or add links. In order to parse wikitext, the community has built libraries such as mwparserfromhell, developed over 10 years and comprising almost 10,000 lines of code. This library provides an easy interface to identify different elements of an article, such as links, templates, or just the plain text. This ecosystem of tooling lowers the technical barriers to working with the dumps because users do not need to know the details of XML or wikitext.
While convenient, there are severe drawbacks to working with the XML dumps containing articles in wikitext. In fact, MediaWiki translates wikitext into HTML which is then displayed to the readers. Thus, some elements contained in the HTML version of the article are not readily available in the wikitext version; for example, due to the use of templates. This means that parsing only wikitext means that researchers might ignore important content which is displayed to readers. For example, a study by Mitrevski et al. found for English Wikipedia that from the 475M internal links in the HTML versions of the articles, only 171M (36%) were present in the wikitext version.
Therefore, it is often desirable to work with HTML versions of the articles instead of using the wikitext versions. Though, in practice this has remained largely impossible for researchers. Using the MediaWiki APIs or scraping Wikipedia directly for the HTML is computationally expensive at scale and discouraged for large projects. Only recently, the Wikimedia Enterprise HTML dumps have been introduced and made publicly available with regular monthly updates so that researchers or anyone else may use them in their work.
However, while the data is available, it still requires lots of technical expertise by researchers, such as how different elements from wikitext get parsed into HTML elements. In order to lower the technical barriers and improve the accessibility of this incredible resource, we released the first version of mwparserfromhtml, a library that makes it easy to parse the HTML content of Wikipedia articles – inspired by the wikitext-oriented mwparserfromhell.
Figure 1. Examples of different types of elements that mwparserfromhtml can extract from an article
The tool is written in Python and available as a pip-installable package. It provides two main functionalities. First, it allows the user to access all articles in the dump files one by one in an iterative fashion. Second, it contains a parser for the individual HTML of the article. Using the Python library beautifulsoup, we can parse the content of the HTML and extract individual elements (see Figure 1 for examples):
Wikilinks (or internal links). These are annotated with additional information about the namespace of the target link or whether it is disambiguation page, redirect, red link, or interwiki link.
External links. We distinguish whether it is named, numbered, or autolinked.
Categories
Templates
References
Media. We capture the type of media (image, audio, or video) as well as the caption and alt text (if applicable).
Plain text of the articles
We also extract some properties of the elements that end users might care about, such as whether each element was originally included in the wikitext version or was transcluded from another page.
Building the tool posed several challenges. First, it remains difficult to systematically test the output of the tool. While we can verify that we are correctly extracting the total number of links in an article, there is no “right” answer for what the plain text of an article should include. For example, should image captions or lists be included? We manually annotated a handful of example articles in English to evaluate the tool’s output, but it is almost certain that we have not captured all possible edge cases. In addition, other language versions of Wikipedia might provide other elements or patterns in the HTML than the tool currently expects. Second, while much of how an article is parsed is handled by the core of MediaWiki and well documented by the Wikimedia Foundation Content Transform Team and the editor community on English Wikipedia, article content can also be altered by wiki-specific Extensions. This includes important features such as citations, and documentation about some of these aspects can be scarce or difficult to track down.
The current version of mwparserfromhtml constitutes a first starting point. There are still many functionalities that we would like to add in the future, such as extracting tables, splitting the plain text into sections and paragraphs, or handing in-line templates used for unit conversion (for example displaying lbs and kg). If you have suggestions for improvements or would like to contribute, please reach out to us on the repository, and file an issue or submit a merge request.
Finally, we want to acknowledge that the project was started as part of an Outreachy internship with the Wikimedia Foundation. We encourage folks to consider mentoring or applying to the Outreachy program as appropriate.
Wikimedia Commons is our open media repository. Like Wikipedia and its other sister projects, Commons runs on the MediaWiki platform. Commons is home to millions of photos, documents, videos, and other multimedia files.
MediaWiki has a built-in imagescaler that, until now, we used in production as well. To improve security isolation, we started an effort in 2015 to develop support in MediaWiki for external media handling services. We choose Thumbor, an open-source thumbnail generation service, for Wikimedia’s thumbnailing needs.
During routine post-deployment checks we found the p99 First Paint metric regressed from 4s to 20s. That’s quite a jump. The median and p75 during the same time period remained constant at their sub-second values.
Distribution of First Paint, which prompted our investigation.
After an investigation we learned that page load time and visual rendering metrics are often skewed in visually hidden browser tabs (such as tabs that are open in the background). The deployment had refactored code such that background tabs could deprioritize more of the rendering work. Rather than revert this, we decided to change how MediaWiki’s Navigation Timing client collects these metrics. We now only sample pageviews in browser tabs that are “visible” from their birth until the page finishes loading.
To understand why background tabs had such an impact on our global metrics, we also ran a simple JS counter for a few days. We found that over a three-day period, 8.4% of page views in capable browsers were visually hidden for at least part of their load time. (Measured using the Page Visibility API, which itself was available on 98% of the sampled pageviews.)
Browser support and distribution of page visibility on Wikipedia.
– Peter Hedenskog and Timo Tijhof.
Performance Inspector goes Beta
We had an idea to improve page load time performance on Wikipedia by providing performance metrics to editors through an in-article modal link (T117411). By using the Performance Inspector, tech-savvy Wikipedians could use this extra data to inform edits that make the article load faster. At least, that was the idea.
It turns out that in reality it’s hard for users to distinguish between costs due to the article content and costs of our own software features. It was hard for editors to actually do something that made a noticeable difference in page load time. We discontinued the Performance Inspector in favor of providing more developer-oriented tools.
— Peter Hedenskog.
The discontinued Perf Inspector offered a modal interface to list each bundle with its size in kilobytes.
The “mw.inspect” console utility for calculating bundle sizes.
Hello, HTTP/2!
Deploying HTTP/2 support to the Wikimedia CDN significantly changed how browsers negotiate and transfer data during the page load process. We anticipated a speed-up as part of the transition, and also identified specific opportunities to leverage HTTP/2 in our architecture for even faster page loads.
We also found unexpected regressions in page load performance during the HTTP/2 transition. In Chrome, pageviews using HTTP/2 initially had a slower Time to First Paint experience when compared to the previous HTTP/1 stack. We wrote about this in HTTP/2 performance revisited.
– Timo Tijhof and Peter Hedenskog.
Stylesheet-aware dependency tracking
2016 saw a new state-tracking mechanism for stylesheets in ResourceLoader (Wikipedia’s JS/CSS delivery system). The HTML we send from MediaWiki to the browser, references a bundle of stylesheets. The server now also transmits a small metadata blob alongside that HTML, which provides the JS client with information about those stylesheets. On the client side, we utilize this new metadata to act as if those stylesheets were already imported by the client.
Why now
MediaWiki is built with semantic HTML and standardized CSS classes in both PHP-rendered and client-rendered elements alike. The server is responsible for loading the current skin stylesheets. We generally do not declare an explicit dependency from a JS feature to a specific skin stylesheet. This is by design, and allows us to separate concerns and give each skin control over how to style these elements.
The adoption of OOUI (our in-house UI framework that renders natively in both PHP and JavaScript), got to a point where an increasing number of features needed to load OOUI both as stylesheet for server-rendered elements, but also potentially load OOUI for (unrelated) JS functionality such as modal interactions elsewhere on the page. These JS-based interactions can happen on any page, including on pages that don’t embed OOUI elements server-side. Thus the OOUI module must include stylesheets in this bundle. This would have caused the stylesheet to sometimes download twice. We worked around this issue for OOUI, through a boolean signal from the server to the JS client (in the HTML head). The signal indicates whether OOUI styles were already referenced (change 267794).
Outcome
We turned our workaround into a small general-purpose mechanism built-in to ResourceLoader. It works transparently to developers, and is automatically applied to all stylesheets.
This enabled wider adoption of OOUI, and also applied the optimization to other reusable stylesheets in the wider MediaWiki ecosystem (such as for Gadgets). It also facilitates easy creation of multiple distinct OOUI bundles without developers having to manually track each with a boolean signal.
This tiny capability took only a few lines of code to implement, but brought huge bandwidth savings; both through relative improvements as well as through what we prevented from being incurred in the future.
Despite being small in code, we did plan for a multi-month migration (T92459). Over the years, some teams had begun to rely on a subtle bug in the old behavior. It was previously permitted to load a JavaScript bundle through a static stylesheet link. This wasn’t an intended feature of ResourceLoader, and would load only the stylesheet portion of the bundle. Their components would then load the same JS bundle a second time from the client-side, disregarding the fact that it downloaded CSS twice. We found that the reason some teams did this was to avoid a FOUC (first load the CSS for the server-rendered elements, then load the module in its entirety for client-side enhancements). In most cases, we mitigated this by splitting the module in question in two: a reusable stylesheet and a pure JS payload.
– Timo Tijhof.
One step closer to Multi-DC
Prior to 2015, numerous MediaWiki extensions treated Memcache (erroneously) as a linearizable “black box”. A box that could be written to in a naive way. This approach, while somewhat intuitive, was based on dated and unrealistic assumptions:
That cache servers are always reachable for updates.
That transactions for database writes never fail, time out, or get rolled back later in the same request.
That database servers do not experience replication lag.
That there are no concurrent web requests also writing to the same database or cache in between our database reads.
That application and cache servers reside in a single data center region, with cache reads always reflecting prior writes.
The Flow extension, for example, made these assumptions and experienced anomalies even within our primary data center. The addition of multiple data centers would amplify these anomalies, reminding us to face the reality that these assumptions were not true.
Flow became among the first to adopt WANCache, a new developer-friendly interface we built for Memcached, specifically to offer high resiliency when operating at Wikipedia scale.
Replication lag was especially important. In MySQL/MariaDB, database reads can enjoy an “isolation level” that offers session consistency with repeatable reads. MediaWiki implements this by wrapping queries from a web request in one transaction. This means web requests will interact with one consistent and internally stable point-in-time state of the database. For example, this ensures foreign keys reliably resolve to related rows, even when queried later in the same request. However, it also means these queries perceive more replication lag.
WANCache is built using the “cache aside” and “purge” strategies. This means callers let go of the fine-grained control of (problematically) directly writing cache values. In exchange, they enjoy the simplicity of only declaring a cache key and a closure that computes the value. Optionally, they can send a “purge” notification to invalidate a cache key during a (soon-to-be-committed) database write.
Instead of proactively writing new values to both the database and the cache, WANCache lets subsequent HTTP requests fill the cache on-demand from a local DB replica. During the database write, we merely purge relevant cache keys. This avoids having to wait for, and incur load on, the primary DB during the critical path of wiki edits and other user actions. WANCache’s tombstone system prevents lagged data from getting (back) into a long-lived cache.
We made numerous improvements to database performance across the platform. This is often in collaboration with SRE and/or with the engineering teams that build atop our platform. We regularly review incident reports, flame graphs, and other metrics; and look for ways to address infra problems at the source, in higher-level components and MediaWiki service classes.
For example, the incident where a partial outage due to database unavailability, was caused by significant network saturation on the Wikimedia Commons database replicas. The saturation occurred due to the PdfHandler service fetching metadata from the database during every thumbnail transformation and every access to the PDF page count. This was mitigated by removing the need for metadata loads from the thumbnail handler, and refactoring the page count to utilize WANCache.
Another time we used our flame graphs to learn one of the top three queries came from WikiModule::preloadTitleInfo. This DB query uses batching to improve latency, and would traditionally be difficult to cache due to variable keys that each relate to part of a large dataset. We applied WANCache to WikiModule and used the “checkKeys” feature to facilitate easy cache invalidation of a large category of cache keys, through a single operation; without need for any propagation or tracking.
Creating a Docker image for your service should be easy—cram your code and its dependencies into a container: boom. done.
But that’s never the whole story.
You have to build new images for each release, monitor them for vulnerabilities, and find a way to safely ship them to production.
You need a reliable process to create, test, and deploy images to Kubernetes. In short: you need a release pipeline.
Wikimedia’s service release pipeline 🚢
A “build and deployment expert” is an antipattern.
Jez Humble & David Farley, Continuous Delivery
Wikimedia has a little more than thirty microservices running atop our in-house Kubernetes infrastructure.
Back when we started moving to Kubernetes in 2017, we had a few aims:
Build trust – After you generate an image, build confidence through incremental testing and validation.
Streamlined image builds – Developer teams shouldn’t need to be experts to build an excellent image for their service.
Security – Build on known-good images, run as a non-root user, and monitor for common vulnerabilities and exposures (CVEs).
And we created two tools to help us achieve these goals:
Blubber – This tool ensures our Docker images are lean, safe, and built from our blessed subset of known-good base images.
PipelineLib – A Jenkins library that uses Blubber to produce, test, and promote images to our Docker registry after establishing trust.
But our migration from Jenkins to GitLab has required some changes to these tools.
Kokkuri: the pipeline from GitLab 🦊
Now we’re migrating to GitLab, we’re replacing Jenkins and PipelineLib with a shared GitLab repository called Kokkuri.
What PipelineLib was for Jenkins, Kokkuri is for GitLab. You can extend Kokkuri jobs in your GitLab project’s `.gitlab-ci.yml` to build streamlined and secure docker images for Wikimedia production, test them, and push them to our production registry.
We’re using this tooling today for two of our internal projects: Scap (our deployment tool for MediaWiki) and Blubber itself.
For now, Kokkuri is an internal tool for Wikimedia’s GitLab. Using it outside of our unique production environment wouldn’t make sense.
Blubber as a BuildKit Frontend 🐳
All of our Wikimedia production services use Blubber to build their Docker images. Blubber is an active, open project—for use both inside and outside Wikimedia 🎉 And as part of the migration to GitLab, we’ve made improvements.
Blubber used to generate opinionated Dockerfiles—now it’s a full-fledged BuildKit front-end. BuildKit is a project from Moby, the people who make Docker, and it’s now used by Docker itself to create images.
As with all in-progress migrations: we’re still missing some things.
Here’s what we’re working on next for our GitLab move:
Dependency caching – tests will be slow if they need to fetch a lot of dependencies for every run, we’re working on a few solutions and you can follow along on Phabricator.
Visibility – we’re still missing all the nice integrations we have in our old systems
Links between our bug tracker (Phabricator) and GitLab
IRC and Slack notifications—yes, we use both 😅
But why “Kokkuri”? 🦝
Tanukis: a crucial part of our pipeline.
Alright. Let’s unpack the name “kokkuri.”
Fun fact: the GitLab logo may look like a fox, but it’s a tanuki—a totally real racoon/dog/fox-type thing `{{citation-needed}}`.
Tanukis are the real-life inspiration for a mythical trickster known as a “kokkuri-san”—an animal spirit bringing mischief, magic, and luck.
And to summon a kokkuri-san: you’d use a kokkuri—which is kinda like a Japanese Ouija board.
So.
To summon a mischievous and magical tanuki you use a kokkuri. And now you can summon our tricksy GitLab magic in the exact. same. way.
I’m happy to share that the second Web Perf Hero award of 2022 goes to Valentín Gutierrez!
This award is in recognition of Valentín’s work on the Wikimedia CDN over the past three months. In particular, Valentín dove deep into Apache Traffic Server. We use ATS as the second layer in our HTTP Caching strategy for MediaWiki. (The first layer is powered by Varnish.)
Cache miss
Valentín (@Vgutierrez) observed that ATS was treating many web requests as cache misses, despite holding a seemingly matching entry in the cache. To understand why, we have to talk about the Vary header.
If a page is served the same way to everyone, it can be cached under its URL and served as such to anyone navigating to that same URL. This is nearly true for us from a statistical viewpoint, except that we have editors with logged-in sessions, whose pageviews must bypass the CDN and its static HTML caches. In HTTP terminology, we say that MediaWiki server responses “vary” by cookies. Two clients with different cookies may get a different response. Two clients with the same cookies, or with no cookies, can enjoy the same cached response. But, log-in sessions aren’t the only cookies in town! For example, our privacy-conscious device counting metric also utilizes a cookie (“WMF-Last-Access”). It is a very low entropy cookie, but a cookie nonetheless. We also optionally use cookies for fundraising localisation, and various other JavaScript features. As such, a majority of connecting browsers will have at least one cookie.
The HTTP specification says that when a response for a URL varies by the value of a header (in our case, the Cookies header controls whether you’re logged-in), then cache proxies like ATS and Varnish must not re-use a cache entry, unless the original and current browser have the exact same cookies. For the cache to be effective, though, we must pay attention to the session cookie only, and ignore cookies related to metrics and JavaScript. For our Varnish cache, we do exactly that (through custom VCL code), but we never did this for ATS.
And so work began to implement Lua code for ATS to identify session cookies, and treat all other cookies as if they don’t exist — but only within the context of finding a match in the cache, restoring them right after.
In our Singapore data center, our ATS latency improved by 25% at the p75, e.g. from 475ms down to 350ms compared to the same time and day a week earlier. That’s a 125ms drop, which is one of the biggest reductions we’ve ever documented!
The reduction is due to more requests being served directly from the cache, instead of generating new pageviews for each combination of unrelated cookies. We can also measure this as a ratio between cache hits and cache misses — the cache hit ratio. For the Amsterdam data center, ATS cache hits went from ~600/s to 1200/s. As a percentage of all backend traffic, that’s from 2% to 4%. (The CDN frontend enjoys a cache hit ratio of 90-99% depending on entrypoint.)
Disk reads
In September, Valentín created a Grafana dashboard to explore metrics from internal operations within ATS. This is part of on-going work to establish a high-level SLO for ATS. ATS reads from disk as part of serving a cache hit. Valentín discovered that disk reads were regularly taking up to a whole second.
Most traffic passing through ATS is a cache miss, where we respond within 300ms at the p75 (latency shown earlier). For the subset where we serve a cache hit at the ATS layer, we generally respond within ~5ms, magnitudes faster. When we observed a cache hit taking 1000ms to respond, that is not only very slow, it is also notably slower than generating a fresh page from a MediaWiki server.
After ruling out timeout-related causes, Valentín traced the issue to the ATS cache_dir_sync operation. This operation synchronizes metadata about cache entries to disk, and runs once every few minutes. It takes about one minute, during which we consistently saw 0.1% of requests experience the delay. Cache reads had to wait for a safety lock held by a single sync for the entire server. Valentín worked around the issue by partitioning the cache into multiple volumes, with the sync (and its lock) applying only to a portion of the data. These are held for a shorter period of time, and less likely to overlap with a cache read in the first place. (our investigation, upstream issue)
On most ATS servers, the cache read p999 dropped from spiking at 1000ms down to a steady 1ms. That’s a 1000X reduction!
Note that this issue was not observable through the 75th percentile measure, because each minute affected a different 0.1% of requests, despite happening consistently throughout the day. This is why we don’t recommend p75 for backend objectives. Left continuously, much more than 0.1% of clients would experience the issue. Resolving this avoids a constant spending of the error budget SLO, preserving our budget for more unusual and unforeseen issues down the line.
Web Perf Hero award
The Web Perf Hero award is given to individuals who have gone above and beyond to improve the web performance of Wikimedia projects. The initiative is led by the Performance Team and started mid-2020. It is awarded quarterly and takes the form of a Phabricator badge.
In 2016, the Wikimedia Foundation deployed HTTP/2 (or “H2”) support to our CDN. At the time, we used Nginx- for TLS termination and two layers of Varnish for caching. We anticipated a possible speed-up as part of the transition, and also identified opportunities to leverage H2 in our architecture.
The HTTP/2 protocol was standardized through the IETF, with Google Chrome shipping support for the experimental SPDY protocol ahead of the standard. Brandon Black (SRE Traffic) led the deployment and had to make a choice between SPDY and H2. We launched with SPDY in 2015, as H2 support was still lacking in many browsers, and Nginx did not support having both. By May 2016, browser support had picked up and we switched to H2.
Goodbye domain sharding?
You can benefit more from HTTP/2 through domain consolidation. The following improvements were achieved by effectively undoing domain sharding:
Faster delivery of static CSS/JS assets. We changed ResourceLoader to no longer use a dedicated cookieless domain (“bits.wikimedia.org”), and folded our asset entrypoint back into the MediaWiki platform for faster requests local to a given wiki domain name (T107430).
Speed up mobile page loads, specifically mobile-device “m-dot” redirects. We consolidated the canonical and mobile domains behind the scenes, through DNS. This allows the browser to reuse and carry the same HTTP/2 connection over a cross-domain redirect (T124482).
Faster Geo service and faster localized fundraising banner rendering. The Geo service was moved from geiplookup.wikimedia.org to /geoiplookup on each wiki. The service was later removed entirely, in favor of an even faster zero-roundtrip solution (0-RTT): An edge-injected cookie within the Wikimedia CDN (T100902, patch). This transfers the information directly alongside the pageview without the delay of a JavaScript payload requesting it after the fact.
Could HTTP/2 be slower than HTTP/1?
During the SPDY experiment, Peter Hedenskog noticed early on that SPDY and HTTP/2 have a very real risk of being slower than HTTP/1. We observed this through our synthetic testing infrastructure.
In HTTP/1, all resources are considered equal. When your browser navigates to an article, it creates a dedicated connection and starts downloading HTML from the server. The browser streams, parses, and renders in real-time as each chunk arrives. The browser creates additional connections to fetch stylesheets and images when it encounters references to them. For a typical article, MediaWiki’s stylesheets are notably smaller than the body content. This means, despite naturally being discovered from within (and thus after the start of) the HTML download, the CSS download generally finishes first, while chunks from the HTML continue to trickle in. This is good, because it means we can achieve the First Paint and Visually Complete milestones (above-the-fold) on page views before the HTML has fully downloaded in the background.
Page load over HTTP/1.
In HTTP/2, the browser assigns a bandwidth priority to each resource, and resources share a single connection. This is different from HTTP/1, where each resource has its own connection, with lower-level networks and routers dividing their bandwidth equally as two seemingly unrelated connections. During the time where HTML and CSS downloads overlap, HTTP/1 connections each enjoyed about half the available bandwidth. This was enough for the CSS to slip through without any apparent delay. With HTTP/2, we observed that Chrome was not getting any CSS response until after the HTML was mostly done.
Page load over SPDY.
This HTTP/2 feature can solve a similar issue in reverse. If a webpage suffers from large amounts of JavaScript code and below-the-fold images being downloaded during the page load, under HTTP1 those low-priority resources would compete for bandwidth and starve the critical HTML and CSS downloads. The HTTP/2 priority system allows the browser and server to agree, and give more bandwidth to the important resources first. A bug in Chrome caused CSS to effectively have a lower priority relative to HTML (chromium #586938).
First paint regression correlated with SPDY rollout. (Ori Livneh, T96848#2199791)
We confirmed the hypothesis by disabling SPDY support on the Wikimedia CDN for a week (T125979). After Chrome resolved the bug, we transitioned from SPDY to HTTP/2 (T166129, T193221). This transition saw improvements both to how web browsers give signals to the server, and the way Nginx handled those signals.
As it stands today, page load time is overall faster on HTTP/2, and the CSS once again often finishes before the HTML. Thus, we achieve the same great early First Paint and Visually Complete milestones that we were used to from HTTP/1. But, we do still see edge cases where HTTP/2 is sometimes not able to re-negotiate priorities quick enough, causing CSS to needlessly be held back by HTML chunks that have already filled up the network pipes for that connection (chromium #849106, still unresolved as of this writing).
Lessons learned
These difficulties in controlling bandwidth prioritization taught us that domain consolidation isn’t a cure-all. We decided to keep operating our thumbnail service at upload.wikimedia.org through a dedicated IP and thus a dedicated connection, for now (T116132).
Browsers may reuse connections for multiple domains if an existing HTTPS connection carries a TLS certificate that includes the other domain in its SNI information, even when this connection is for a domain that corresponds to a different IP address in DNS. Under certain conditions, this can lead to a surprising HTTP 404 error (T207340, mozilla #1363451, mozilla #1222136). Emanuele Rocca from SRE Traffic Team mitigated this by implementing HTTP 421 response codes in compliance with the spec. This way, visitors affected by non-compliant browsers and middleware will automatically recover and reconnect accordingly.