Normal view

There are new articles available, click to refresh the page.
Before yesterdayMain stream

Wikimania 2026: Growing language communities

24 July 2026 at 18:28

Are large language models widening the digital divide between majority and minority languages? Or can they be harnessed in preserving language diversity?

On Day 3 of Wikimania Paris, representatives of many different language and cultural communities shared their experiences in working within the Wikimedia movement.

Respect for indigenous communities

Erroneous information about indigenous languages and cultures can alienate potential contributors. During the keynote panel, Michelle Collipal, a member of the indigenous Mapuche community in Chile, shared how she has started to motivate Mapudungùn speakers and cultural authorities to engage with the Mapudungùn Wiki project. 

Moderator Dr Terri Janke underscored the importance of involving indigenous communities, respecting their knowledge as well as their right to self-determination. What this means in practice for wiki spaces is discussed in the white paper “CutureStrong Platforms: Setting the Standard at Wikimedia”.

Automating and improving translation

During “Language Diversity in the Digital Commons”, Pau Giner of the Wikimedia Foundation shared that the MinT (Machine in Translation) was designed from the start to support diverse language models. The platform currently supports 200 languages.

His ideas for the future include making Wikipedia articles more accessible on mobile phones, and to offer a collection of tools to help people create their first wiki, going beyond translation.

During the Q&A, Kepa Sarasola (User:karasola) of University of Basque reflected on the improvement of translation tools available since he started working on Basque Wikipedia in 2011. The tools, he said, had made it possible to grow minority language Wikipedias – while giving them the flexibility to choose which parts of articles to translate and where to start from scratch. The integration of AI tools has accelerated the process significantly.

Collaboration across communities

The need for collaboration across language communities was a recurring theme across multiple sessions. One example of a cross-border network is Linguatec-IA, an EU project for the digitization of languages in communities in the Pyrenees region, including Basque, Catalan, Occitan, and Aragonese.

David Castillo Parra of UNESCO discussed the launch of the New Commons Incubator for indigenous-led capacity building programs. Applications will be accepted for indigenous-led teams through 14 August.

The need for the Wikimedia movement to remain true to its mission – and to aligning on milestones that matter – was an inspiring message from Audrey Tang in “The State of Wikimedia & AI 2026”. “Don’t let it be a race…Wikipedia never tried to win. We tried to make sure there were many winners at any one time.”

As Jimmy Wales told Le Monde, “It’ll be all right. We’ll adapt, change, use AI in our own way. We’ll find a way.”

Wikimania 2026: Freedom, equity and reliability

23 July 2026 at 17:21

How can the Wikimedia community defend freedom, equity, and reliability on the internet? Are we in a global information crisis?

Several sessions on Day 2 of Wikimania Paris explored these questions from different angles. The morning kicked off with breakout sessions on how global trends are impacting government regulation – with attendees joining Wikimedia Foundation board members in small group discussions. Protecting free knowledge was a core theme throughout the day.

Misinformation on climate change

During the 2025 Iberian peninsula blackout, misinformation was spread that the main cause was renewable energy. In “Climate Conversations: Interdisciplinary Approaches to Knowledge Sharing”, moderator Tatjana Baleta said that rapid response editing on Wikipedia was one way that the scientific community combats misinformation during extreme weather events and other major incidents.

Many readers have also shifted from reading about climate change to reading more about other issues such as the cost of living crisis. Dr. Femke Nijsse (User:Femke) discussed the importance of meeting readers where they are – and explaining the science related to these issues.

While AI-generated content may have the sheen of reliability, it often turns out that the sources they cite are hallucinations or that they do not actually verify the claims made by the LLM. Editors on Wikipedia are now starting to use tools such as AI Source Verification (which itself uses LLMs) to predict the verifiability of claims made in a article.

AI crawlers: Encroaching on creativity?

“Collateral Damage? Human Creativity and Interaction in the AI Crawling Era” explored how organizations in the free knowledge ecosystem are responding to the massive increase in AI scrapers.

There was a consensus among panelists that attribution is a critical concern for authors. Creative Commons CEO Anna Turnadóttir and Monica Westin of Cambridge University Press discussed the need to educate authors about the benefits of open access models – while also acknowledging the need to experiment.

Mark Graham discussed how the Internet Archive is reaching out to news organizations to discuss alternatives to blocking the Wayback Machine, such as rate limiting and allowing access only for certain uses.

Striking a balance

Throughout the day, speakers debated the complexities of regulation – and how the rush to “do something” can backfire. Turndóttir said, “What used to be the internet handshake online is now the middle finger. And that sort of environment forces lawmakers to reach for blunt tools like regulation. The community needs to establish norms, but sometimes regulation does cause real harm.” 

The keynote session on “Protecting Free Knowledge – The New Battlegrounds of Digital Freedom”, Nnenna Nwakanma (from the internet) discussed the discourse around regulation – and how European approaches may not be applicable in Africa. Panelists during the session, including Axelle Lemaire, architect of the 2016 loi numerique, emphasized the need for balance in protecting openness and freedom online, while also protecting the privacy of individuals.

The conference is making waves in France with media coverage in 25 outlets – including La Croix‘s print edition, a Radio France podcast, and more.

Indic Wikimedia Hackathon Hyderabad 2026

23 July 2026 at 16:00
Group photo of the Indic Wikimedia Hackathon Hyderabad 2026, Image by Nivas

Program Purpose

Fifty-six contributors gathered at IIIT Hyderabad for three days to improve Wikimedia’s technical ecosystem. Unlike traditional hackathons that focus primarily on rapid prototyping, the Indic Wikimedia Hackathon 2026 introduced dedicated refinement sessions that encouraged participants to improve code quality, documentation, usability, and long-term maintainability.

The Indic Wikimedia Hackathon Hyderabad 2026 was organized by Indic MediaWiki Developers User Group (aka Indic-TechCom). The hackathon took place in Hyderabad from 26 – 28 June 2026 (with 25 June as Day 0), in collaboration with the Open Knowledge Initiatives team and OSDG club at International Institute of Information Technology, Hyderabad.

Wikimedia hackathons are spaces for developers, designers, content editors, and other community stakeholders to collaborate on building technical solutions that help improve tools, workflows, and overall user experience across Wikimedia projects.

This hackathon is designed for:

  • Technical contributors active in the Wikimedia technical ecosystem, which includes developers, maintainers (admins/interface admins), translators, designers, researchers, documentation writers, etc.
  • Content contributors having an in-depth understanding of technical issues in their Wikimedia projects, like Wikipedia, Wikisource, Wiktionary, etc.
  • Contributors to any other open-source community or those who have participated in Wikimedia events in the past, and would like to get started with contributing to Wikimedia technical spaces.

Participants worked on a curated set of technical tasks prepared by mentors and organizers. They were also encouraged to propose their own project ideas, provided they included a clear problem statement, implementation approach, and were reviewed by mentors before the event. 

Building on the experience and learnings from previous hackathons, this event was more efficient, inclusive, and collaborative. 

The event aimed to involve more developers who have experience with the Wikimedia ecosystem and had prior experience already contributing to tools, extensions, gadgets, or other technical projects. Editors were paired up with developers to provide domain knowledge, helping them better understand editing workflows, user needs, and the intended behaviour of the applications/ extensions/ gadgets being developed. 

Unlike other hackathons where rapid development is the primary focus, this event was not completely hacking but also incorporated dedicated  refinement sessions.These sessions encouraged participants to improve the quality of the works by refining  design, UI, data privacy, code optimization, documentation and overall maintainability. Additionally, the program also included brainstorming sessions, group discussions, workshops  to help participants  understand a broader perspective of this ecosystem beyond their individual projects.

Scope and Timeline

The hackathon was conducted as a three day in-person event  at the International Institute of Information Technology, Hyderabad (IIIT-H), a long-standing partner that provides space for technical and community events. The venue supported collaborative work through dedicated hacking spaces, mentor interactions, and discussion areas. 

The scope of the event was flexible enough to encourage participants to work on a curated set of Wikimedia-related technical projects and tasks suitable for a hackathon which prepared by mentors or a custom project proposed by their own and reviewed by experienced developers and organizers, along with Team Challenges from Wikimania Hackathon 2026.

The program was structured into distinct phases. The first half of the event (approximately one and a half days) focused entirely on development where the first part of the first day was catered to some introductions and welcome notes, ground rules, ice breaker activities, followed by continuous hacking. On the second day, the event had a social activity and the morning session was mostly catered to workshops, brainstorming sessions, and getting to know about OKI work. The second half started with refinement phases. The third day focused completely on refinement, wrap-up and showcase. 

Attendance

Total Attendees: 56

  • Organisers:   11
  • Mentors:   15
  • Participants:  25
  • Editors:  5

The majority of participants were developers with prior experience in the Wikimedia technical ecosystem. A smaller group consisted of experienced Wikimedia editors with technical knowledge, who collaborated with developers by providing domain expertise and user perspectives during the hackathon. 

Activities Conducted

Prior to the hackathon, an orientation call was conducted for participants to give an overview of the program, the Wikimedia technical ecosystem, team formation details, and some logistical and operational arrangements. The session also introduced participants to Wikimedia, its technical ecosystem, guided them through basic account setup, and explained the overall hackathon format. 

Following the orientation,  participants were encouraged to have a call with their specific teams and mentors. These discussions helped participants better understand their assigned projects, including the scope, objectives, expected outcomes, and technical requirements. Mentors introduced project-specific workflows, outlining the scope, objectives, and tasks for each participant to help them engage effectively during the hackathon.

During the event, participants worked on pre-curated tasks across multiple Wikimedia-related projects and collaborated closely with mentors to understand issue tracking, patch submission, and debugging workflows. Mentors supported participants across different projects, helping them navigate both technical challenges and Wikimedia-specific contribution processes.

To complement the technical program, the hackathon also included community-building activities. An icebreaker session at the beginning of the event helped participants interact and build connections. On the second day, interested participants joined a social walk around Hyderabad, providing an informal opportunity for networking. A dedicated women’s dinner was also organized to foster stronger connections among women participants, encourage inclusion, and support long-term retention within the Wikimedia technical community. 

Towards the conclusion of the event, a project showcase and presentations session was conducted, during which participants demonstrated their work and shared learnings with fellow participants, mentors, and organizers.  

Outputs and Outcomes

The hackathon enabled participants to work on pre-identified tasks across multiple Wikimedia-related repositories, resulting in code contributions, feature enhancements, bug fixes, and documentation improvements. Throughout the event, participants gained practical experience with Wikimedia development workflows, including issue tracking, patch submission, code review, and collaborative problem-solving, with guidance from experienced mentors. Several participants continued engaging with their assigned projects after the event, indicating effective onboarding into Wikimedia technical workflows.

During the hackathon, participants submitted a total of 56 Phabricator tickets across multiple Wikimedia-related projects. The distribution of contributions is summarised below:

  • Clip2Commons:  9
  • Deployr: 7
  • Language Selector Rewrite: 1
  • Lingua Libre: 5
  • Montage:  4
  • NPOV Drift Detector: 1
  • Observability Tool: 7
  • Onboarding of New Wikipedia Editors  : 2
  • OpenSpeaks Subtitler: 5 
  • OpenSpeaks Tome: 3
  • Pywikibot  : 3
  • Scribe: 6
  • Translate Tagger: 8
  • ULS Extension :8
  • Wanda Extension: 4
  • WikiEval Tool : 3
  • Wikievol Tool  : 3
  • WikiLinkua  : 2
  • Wikimedia Commons Android: 4 
  • Wikisource Reader App: 2

Total: 87 Repo/Phabricator tickets linked

As part of the Indic Wikimedia Hackathon Hyderabad 2026, several teams worked on projects that directly align with the official Team Challenges announced for Wikimania 2026. 

  • OpenSpeaks (Subtitler/Tome/Bento) – Boost multimedia experience, Connect multilingual knowledge
  • Lingua Libre – Boost multimedia experience, Connect multilingual knowledge
  • WikiLinkua – Gamify knowledge, Stream data with Wikidata
  • Wiki Translate Tagger – Connect multilingual knowledge
  • Wanda / WandaScore / WandaScribe – Welcoming newcomers, Fix the sources / Update the obsolete, The editor of the future
  • Scribe – Stream data with Wikidata, Connect multilingual knowledge
  • WikiEvolution – Explore knowledge
  • WikiNPOV Drift Detector – Deciphering biases
  • Language selector rewrite – Connect multilingual knowledge

These contributions included code changes, improvements, and related updates submitted under mentor guidance.

What went well:

Program design:

  • The overall event design enabled participants to collaborate with peers from diverse backgrounds and work effectively on technical projects.
  • The combination of structured onboarding, continuous hacking, and dedicated refinement sessions supported steady progress throughout the event.
  • Refinement sessions encouraged participants to improve code quality, documentation, user interface, and maintainability rather than focusing solely on completing tasks.

Mentorship and technical contributions:

  • Clear project introductions and continuous mentor support helped participants engage confidently with their assigned tasks.
  • Most of the identified hackathon tasks were actively worked on during the event.
  • Several participants continued contributing to their assigned projects after the hackathon, demonstrating successful onboarding into Wikimedia technical workflows.

Collaboration:

  • Pairing editors with developers proved valuable, as editors helped developers better understand user workflows, expected tool behaviour, and usability considerations.
  • Workshops and discussion sessions complemented the hacking sessions by providing participants with a broader understanding of the Wikimedia technical ecosystem.

Diversity and inclusion:

  • The hackathon achieved approximately <>% women participation, the highest among events organized by the User Group to date.
  • The women’s dinner helped foster stronger connections among women participants and contributed to a more welcoming environment.

What can be improved/Learnings?

Program schedule

  • Participants requested additional time for the project showcase and presentations.
  • Fifteen-minute breaks were considered too short and could be extended in future editions.
  • Starting sessions at 9:00 AM posed challenges for some participants because of commuting time.

Project selection

  • Hackathon tasks require more review before the event to reduce duplication of effort.
  • Greater emphasis should be placed on improving existing Wikimedia tools rather than developing new ones where similar solutions already exist.

Mentorship

  • Participants experienced delays when waiting for scheduled online mentor support.
  • Increasing mentor availability or ensuring more mentors are physically present during the event could improve the overall experience.

What’s next:

Based on the outcomes and observations from the Indic Wikimedia Hackathon Hyderabad 2026, the following recommendations are proposed for future editions of the event:

  • Prioritise improving existing Wikimedia tools and applications over developing new ones, where appropriate.
  • Continue incorporating dedicated refinement sessions to improve the quality and sustainability of contributions.
  • Allocate more time for project showcases, presentations, and participant discussions
  • Review the event schedule by extending break durations and considering a later start time where feasible.
  • Continue initiatives that promote diversity and inclusion, including activities that support participation and retention of women contributors.
  • Strengthen post-event follow-up and mentorship to encourage continued contributions beyond the hackathon.
  • Explore organizing additional hackathons, workshops, and technical events in college campuses to improve accessibility and encourage new technical contributors.

From a Wild Idea to Reality: Building WISE at Wikimedia Hackathon 2026

23 July 2026 at 11:00
Wikimedia_Hackathon_2026_group_photo_03
Wikimedia_Hackathon_2026_group_photo_03

When people think about hackathons, they often imagine coding sessions, project demos, and late nights spent debugging. For me, Wikimedia Hackathon 2026 in Milan was something much more meaningful: a reminder that some of the most impactful ideas begin as simple conversations between passionate people.

Like many Wikimedia Commons contributors, I have often found it difficult to discover media through traditional search. Commons hosts millions of images, videos, and audio files, but today’s search relies mostly on filenames, categories, descriptions, and structured data not on what is actually visible or audible in the file itself. Hundreds of campaigns, contests and content initiatives add new media every year, which only makes the problem bigger. I kept wondering, half as a joke and half seriously: what if we could search Commons based on what an image or video actually shows, rather than what someone happened to type into its metadata?

It felt like a wild idea at the time, so I shared it with the Wikimedia technical community mostly to see if anyone else found it interesting.

Eugene, David and Gopa at WMHACK 2026
Eugene, David and Gopa at WMHACK 2026

A researcher and developer named David, who had recently joined the Wikimedia community, came across the discussion. David had been working on a technology called WISE.. at the University of Oxford research focused on semantic understanding of visual and multimedia content, enabling people to search images, videos, and audio using natural language. What started as a random online exchange quickly turned into a collaboration. David hadn’t originally planned to attend the hackathon in Milan, but as we kept talking, we both got excited about bringing this kind of search to the Wikimedia ecosystem, and decided to meet in person and try to build it.

Looking back, that decision changed everything.

For two intense days at the hackathon, we worked side by side to integrate and demonstrate WISE for Wikimedia Commons — architecture, datasets, search quality, the usual string of small technical fires. One moment that stuck with me: the first time we typed “horse in an airplane” into the prototype, half-expecting nothing, and watched it actually return the right video. That was the point it stopped feeling like a demo and started feeling like a real tool.

What inspired me as much as the technology was David’s approach to problem-solving. In an era where most developers reach for an AI assistant the moment something breaks, David would just as often dive into technical documentation, research papers, and manuals first. Watching him methodically work through a problem rather than skip to an answer was a good reminder that strong engineering fundamentals haven’t gone anywhere.

What we built

The result of those two days is WISE, a new experimental search experience for Wikimedia Commons. It currently indexes Media of the Day content around 5,000 videos and is built to understand the actual visual and audio content of a file rather than its metadata.

Project Wise showcasing Semantic Image Search
Project Wise showcasing Semantic Image Search

Semantic search. Search using natural language and find media based on what appears in the image or video itself. A few examples that worked surprisingly well:

  • “man at a train station”
  • “horse in an airplane”
  • “man with a flower”
  • “pirate with a pistol”

Audio search. Search within audio content to find relevant recordings and segments.

Project Wise showcasing Facial Recognition Feature
Project Wise showcasing Facial Recognition Feature

Face search. Upload a photo of a face, and WISE can locate that person across images and videos… for video, it can even surface the timestamps where they appear.

Multilingual search. Queries work in multiple languages, including Hindi and Telugu. This matters more than it might first appear: most of Commons’ existing search tooling is built around English-language metadata, which quietly shuts out a large share of Wikimedia’s global, non-English-speaking contributor and reader base. A search experience that understands a query in Hindi or Telugu as well as it understands one in English is a small but real step toward making Commons more usable for the movement it actually serves.

You can try WISE yourself here: wise.wmcloud.org
Commons Project page: commons.wikimedia.org/wiki/Commons:Wise

What’s next

This is only the beginning. We’re already discussing:

  • Expanding indexing beyond Media of the Day to cover Commons at a much larger scale.
  • Finding visually similar images after an upload.
  • Suggesting categories, filenames, and metadata based on visual similarity.
  • Improving search quality and broadening multilingual support further.

Most of all, this experience reminded me why I love being part of the Wikimedia movement. A passing idea shared on a community forum connected two people from different backgrounds and different parts of the world. An online discussion became an in-person collaboration. A concept became a working prototype. And a hackathon became the place that vision came to life.

Sometimes the most valuable outcome of sharing an idea isn’t the idea itself it’s the people who find it, connect with it and decide to build something together. For me, WISE for Commons is more than a search tool. It’s proof of that.

I built the convertor from Wikipedia dumps to man (roff) format – for offline reading in a terminal

25 June 2026 at 16:00

I love Wikipedia, command-line interface (cli) and respect offline stuff. Finally I did it – Rust software (for peformance), code generated by llm gpt-5.5 xhigh.

How to use: download the Wikipedia dump, and

cargo run --release -- path/to/dump.xml -o path/to/output_directory

Yes for offline reading we have Kiwix – but terminal is also good. For you batery as well. Or when you system in under heavy load so your CPU/RAM are limited. Man format mean that you can read Wikipedia through SSH as well, and from very cheap devices.

This is not perfect – we should add special handlers for different templates. But it already mostly works and valuable for the comunity. Try it, ping if you want to improe something. You are welcome to pack this software for you linux distribution – I packed only for Gentoo.

Wikipedia in man screenshot

See the project repository, with instructions, at https://gitlab.com/vitaly_zdanevich_wikimedia/wiki2man_on_rust

Documenting Nigeria’s Rites and Rituals through Wiki Loves Africa 2026 for Creatives in Nigeria

By: Meritkosy
25 June 2026 at 10:00

For years, rites and rituals have shaped the identity of communities across Nigeria. From naming ceremonies and traditional weddings to religious observances and everyday customs, these practices preserve values, beliefs, and collective memories. Yet many of these cultural expressions remain underrepresented on Wikimedia projects.

 Wiki Loves Africa 2026 for Creatives in selected Nigerian universities broughtt together photographers and storytellers to document and share Nigeria’s rich cultural heritage under the theme “Rites and Rituals.” The campaign was held from February to April 2026.

Creating awareness and recruiting creatives

Between 1 and 28 February 2026, the campaign focused on awareness and participant recruitment. Outreach activities targeted creatives across six institutions and communities, while social media campaigns and sponsored advertisements on Facebook and Instagram helped expand participation.

These efforts introduced creatives to Wikimedia Commons and encouraged them to contribute freely licensed media that celebrate Nigeria’s diverse traditions.

Flyer for Awareness

Launching the campaign and introducing Participants to Wikimedia Commons

The campaign officially kicked off with a virtual launch held on 4 March 2026. The session introduced participants to the Wiki Loves Africa 2026 theme, “Rites and Rituals,” and provided an orientation on contributing to Wikimedia Commons.

Participants learned about:

  • Wikimedia Commons and its role in preserving open knowledge;
  • Creating Wikimedia accounts;
  • Understanding free licenses;
  • Best practices for successful contest submissions.

The virtual launch provided a foundation for participants, which some were first-time contributors to Wikimedia projects.

Training sessions

Throughout March and April, participants joined several training sessions organized by the international Wiki Loves Africa team. These sessions equipped creatives with technical and ethical skills required for documenting cultural heritage responsibly.

Some of the sessions included:

  • The Wiki Loves Africa International Virtual Launch on 5 March 2026;
  • A photo essay writing workshop on 11 March 2026;
  • Audio creation and sound design training for Wikimedia Commons on 19 March 2026;
  • Ethical storytelling training on 27 March 2026;
  • “Documenting our Rites and Rituals with Respect and Integrity” on 28 March 2026;
  • Pattypan mass upload training on 4 April 2026;
  • Uploading photographs, videos, and audio files to Wikimedia Commons on 10 April 2026.

These learning opportunities strengthened participants’ understanding of visual storytelling, copyright, ethical documentation, and the technical processes involved in contributing multimedia content to Wikimedia Commons.

Taking the campaign to communities through photowalks

Beyond virtual engagement, the campaign emphasized practical documentation through physical meetups and photowalks.

Creatives in Rivers State held their meetup and photowalk between 3–5 April 2026, where participants explored their communities and documented the Easter Rites.

On 11 April 2026, creatives in Abuja gathered for their meetup and photowalk, providing opportunities for collaboration, peer learning, and hands-on experience in capturing images related to the theme. Gombe creatives hosted their physical meet-up on 25th April 2026, and so did Edo State and Oyo State.

Physical meet-up of creatives in Abuja
Physical meet-up of creatives in Abuja
During physical meet-up

Every photograph, audio recording, and video uploaded to Wikimedia Commons contributes to a growing repository of freely accessible knowledge, ensuring that Nigeria’s rites and rituals are visible to people around the world and preserved for future generations.

As participants continue contributing beyond the campaign, Wiki Loves Africa remains a reminder that documenting culture is not only about preserving the past. It is also about ensuring that African stories are represented and shared openly with the world.

Submissions can be explored through the event page , the dashboard or through the category

Commons Quality Images, 20 years of celebrating Wikimedians taking those photographs we need

By: Gnangarra
25 June 2026 at 07:00

Commons Quality Images (COM:QI) was created during June 2006. The purpose of COM:QI was to identify high quality photographs taken by Wikimedians and made freely available to the community at large through Commons. As of 14 June 2026 over 450,000 media files, mostly photographs, have been recognised through this process. Alongside photographs, QI has Microscopic images, animated GIF’s, and graphical works including our own QI seal (pictured).

Green wax seal
Commons Quality Image seal

From the Beginning

In September 2005 Commons became live, just 6 months later I was encouraged over by User:Pfctdayelise who had been there from the first month. In June 2006 User:Pfctdayelise was searching to put together collections of images created by Commons users as rotating background images, or calendars to help promote the project. I tried to help, we could not even find a common subject matter consisting of reasonably sized and quality printable images. Wikimedia Commons had some great photos, some of which had been featured, but a substantial portion were images scraped from other sites like NASA and Flickr.

That started me thinking about how we raise the profile of community uploaded images and encourage efforts to improve them. Having already contributed to some Good Articles over on en.Wikipedia I thought we could create a similar project but for photographs.

Creating the project I saw some key necessities: it needed to be efficient with clear standards, while not taking weeks to decide the outcome. Most importantly it needed to focus only on images that had been taken by members of the Commons Community and that once recognised it could not be taken away. From there working with User:Wikimol and others we built the project. We also created image guidelines which would enable consistent reviewing of nominations.

The image guidelines started out as;

  1. Photographic images
    1. Low resolution / thumbnail. Photographic QP has to have at least 1.92 megapixels (=1600×1200, for example)
    2. JPEG problems. Too much compressed, too low JPEG quality settings in camera / when saving. Visible jpeg artifacts. => Use better quality settings (e.g. set JPEG “superfine”, shoot RAW, save in photoshop with max. quality)
    3. Noise problems, too much noise. Be it chroma noise, luminance noise, visible grain, scratches in scans… QP should not have distracting amount of noise when viewed in 100%
    4. Bad exposure. Overexposure, blown out highlights, underexposure, shadows details replaced by jpeg maps… In incorrectly exposed images, significant details in a significant part are lost.
    5. Color problems. Bad white balance. Distracting (typically purple) hazing at 100%. Color aberration. QP must have reasonable colors (which does not necessarily mean natural colors).
    6. Improper or undefined focus, insufficient depth of field. QP should have clearly defined focus, e.g. main subject in focus, foreground and background out of focus. Or the whole scene in focus. Counterexample – main subject blurry, foreground even more blurry, focus is somewhere between main subject and background. DOF could be low on purpose.
    7. Blur. Images blurred just because of shaking hand or subject moving too fast. Motion blur in QP has to have purpose.
    8. Poor lighting. Including: distracting reflections (usual problem with built-in flash), unintended vignetting, distracting harsh shadows. Generally bad lighting makes scenes with space look flat.
    9. Overfiltered. There are so many PS/Gimp filters. Rarely a better image is created just by applying more and more filters…
    10. Bad or nonexistent composition, unclear or nonexistent subject. QP should have subject and composition of the image should support depiction of that subject, not distract from it.
    11. Bad perspective, tilt, and other distortions. An eye (or, more precisely, a brain) is a sensitive detector capable of spotting even a small tilt … falling trees, churches, inclined water surfaces,… Images of architecture should usually be rectilinear and without too much perspective distortion.
  2. Stitched images, panoramas.
    1. Panoramatic QP has to have a height of 800px min.
    2. Stitching problems. Stitched images should be without artifacts, colors and lightness should be the same across the image.

All of these may look familiar to many of you as they are still present as Commons:Image guidelines, the very same guidelines that every WikiLoves or similar competition uses as their rules. 

Now we had a process with guides, we approached User:LadyofHats who was known for the drawings she was uploading at that time. LOH was asked to create a seal which we could use to help identify successful photos, that seal is the one we continue to use. Some time during this the QI seal itself got recognised as a Quality Image.

In those early months every image was reviewed and the pages manually updated, a time consuming process that would occasionally be interspersed with edit conflicts. The community embraced QI and a new project called Valued Image emerged for recognising sets of image rather than individual images. It was the efforts of User:Dschwen, who had also taken on maintenance tasks, decided to create a bot that would do most of the work for us. Mike Peel continues to maintain the QICbot and has been invaluable at keeping it working, as of 16 June 2026 QICbot has performed 922,000 edits looking after QI tasks.

This photograph of a Western Gull (Larus occidentalis) was the first Quality image to be promoted to Featured Picture status by User:Dschwen

QI grew fast once the QICbot came on board. I stepped into the background watching the community grow QI. Over time many contributors have made QI. Like User:Poco a poco who presented at Wikimania in London on how QI helped him improve his contributions. For the curious there is an opt-in unaudited list of QI by photographers. Whether it’s one successful image or 20,000 of them they all make QI what it is. 

The future

I think some of QI’s potential still hasn’t been realised. I saw it as a historical record of the growth of photography, of something researchers could look back over and see how it has matured. Perhaps more could be done to integrate QI images in content on other projects as they are recognised as our better works. Maybe a bot could triage nominations to identify the more regular issues that cause images to be rejected, though the human touch should always remain the final judge. 

In the future, QICbot will reach 1 million edits, QI will reach 500,000 very soon. QI carries within itself a lot of untapped potential. 

Could it be time to consider creating a sister project for Quality Videos? I know one thing: there will always be enough high quality images created by Wikimedians to make a calendar for any subject. The challenge will be in choosing just 12.

I would must take a moment to acknowledge everyone who has already helped Commons Quality Images along it’s 20 year journey there has been so many as well as everyone who joins the efforts in future. It’ll always be nice to reflect and be able to say I was able to open the door to something so special. The future of Quality Images is now a journey the Wikimedia Commons Community will decide. 

QI’s prosperity comes from the collective effort of everyone! Perhaps this will encourage you to join the QI club.

Team Challenges: New Perspectives and Innovation at Wikimania 2026

24 June 2026 at 18:30

For 25 years, the Wikimedia movement has been constantly innovating to promote the values of free and universal sharing of all human knowledge for the benefit of humanity. In 2026, as we face the challenges of connectivity, misinformation, and the rise of AI, our ability to innovate collectively is more essential than ever.

To address these developments, we are introducing a new program focusing on Team Challenges in the lead-up to Wikimania. This year, we are inviting experts and newcomers to the Wikimedia community to join forces with participants in our Wikimania Hackathon. The idea is not to replace the hackathon format, but to offer a different approach. We want to bring together diverse perspectives and skills by pairing Wikimedians from all walks of life (not just developers!) with professionals from other fields. The goal is to share our best practices and design digital tools that are sustainable, inclusive, and accessible to all.

Next week, we are kicking off our first online orientation sessions training newcomers on wiki tools. In the coming weeks, they will form teams with Wikimedia community members to undertake one of the 2026 challenges. Stay tuned for further updates as the teams work together toward the Team Challenges showcase concurrent with Wikimania Paris. 

We can’t wait to see what the cross-pollination between different disciplines, perspectives, and skill sets will bring. Good luck to all!

10 challenges for maximum impact

The goal is simple: To turn the community’s needs into concrete solutions. We have selected 10 major technical challenges that address the complex issues encountered on a daily basis across Wikimedia projects. 

These challenges are based on the three pillars of Wikimania 2026:

  • FREEDOM: Ensuring the interoperability of free software to facilitate its reuse and maintenance, and to guarantee its long-term viability.
  • EQUITY: Promoting multilingualism. Designing tools capable of overcoming language barriers to bridge gaps in sources and contributions on a global scale.
  • RELIABILITY: Developing AI-based detection tools to assist volunteer moderators in their work of verifying and protecting information.
Discover the 10 Challenges
Looking at a blue puzzle piece through a magnifying glassFix the sources Automatically detect outdated, retracted or broken sources in articles, and facilitate their updating or archiving.

Child and older man sitting at a table doing a puzzle together
Welcoming newcomers Reinvent how we onboard newcomers through chatbots, guided micro-edits, and gamification lower the barrier to entry on wikis.
Birthday cake to celebrate the 25th anniversary of WikipediaUpdate the obsolete Identify and flag outdated articles, obsolete illustrations and screenshots, and outdated graphics to: encourage the community to update them.
Wikimania Paris lightbulb character with a megaphoneBoost multimedia experience Enhance the audiovisual experience on wikis: playlists on Commons, a modernised video player, annotation of audio segments, and games focused on accents and pronunciations.
Statue of a thinker with a pigeon on his headDeciphering biases Develop tools to detect gender, geographic, and cultural biases in Wikimedia content to: encourage the community to correct them.
Wikimania Paris lightbulb character kicking a puzzle globe footballGamify knowledge Create serious games and playful experiences using Wikimedia data to: learn, contribute, and explore in new ways.
French waiter with a serving tray of books and knowledgeStream data with Wikidata Take advantage of Wikidata content to: generate lists, tables or up-to-date visualisations, semi-automatically, for other projects in the Wikimedia galaxy or further.
Laptop with Wikimania Paris mascot Luce the lightbulb visible on the screenThe editor of the future Improve the editing experience: search & replace in the source editor, a standardized citation toolbar, a better visual table editor, and text alignment in VisualEditor to modernise and streamline contributions.
Explore knowledge Design and experiment new ways of sharing knowledge, in interactive and visual ways…for the modern web!
Connect multilingual knowledge Facilitate translation, the reuse of content across Wikimedia projects, and collaboration with other Open Source projects so that knowledge can flow freely in all languages.

For more information on the Team Challenges and how to participate, please visit the Team Challenges section of the Wikimania wiki. Please note that registration for the Team Challenges does not grant access to the rest of Wikimania.

Tech News 2026 – Issue 26

24 June 2026 at 13:00

Latest tech news from the Wikimedia technical community. Please tell other users about these changes. Not all changes will affect you. Translations are available: 日本語, Bahasa Indonesia, Bahasa Melayu, Deutsch, Tiếng Việt, español, français, polski, português, português do Brasil, čeština, русский, українська, עברית, العربية, 中文

Weekly highlight

  • Growth features are now available at Wikidata. This update enables access to Mentorship (if configured), Impact module, the Help Panel, and a simplified Newcomer Homepage (without Suggested Edits). Wikidata administrators are still configuring the features through Community Configuration.

Updates for editors

  • The special page Special:RangeCalculator has been created. It allows users to find an IP range without needing to rely on external tools. Until now, this tool was only available to CheckUsers. [1]
  • Sub-referencing is a new MediaWiki feature that allows editors to reuse references with different details. It will be deployed to most small and medium-sized Wikipedia language versions on June 23. The FAQ lists possible actions to take on your wiki to support the deployment. Check the rollout plan for the next deployment steps. [2]
  • Starting next week, users will get a notification when they are blocked or unblocked from editing, or if this block changes. [3]
  • Recurrent item View all 32 community-submitted tasks that were resolved last week.

Updates for technical contributors

  • Starting next week, abuse filters that are set to “require CAPTCHA verification” will begin to also affect users with the skipcaptcha right, which includes most autoconfirmed users. Bots are exempted. This change only affects edits that trigger an abuse filter. The skipcaptcha right will continue to exempt users from having to solve CAPTCHAs in the ordinary course of using the wikis. [4]
  • Reference documentation for the Lift Wing API has moved from the API Portal to the interactive REST Sandbox.
  • The API Portal wiki is now closed. For API documentation, see Wikimedia APIs on mediawiki.org. All API Portal wiki URLs (https://api.wikimedia.org/wiki/) will redirect to the mediawiki.org page starting June 22. [5]
  • Recurrent item Detailed code updates later this week: MediaWiki

Meetings and events

Tech news prepared by Tech News writers and posted by bot • Contribute • Translate • Get help • Give feedback • Subscribe or unsubscribe.

Wikis for Everyone: Bridging the Accessibility Gap at the 2026 Hackathon

2 June 2026 at 13:00
Wikimedians discussing web accessibility at the Wikimedia Hackathon 2026
Italian wikimedians discussing web accessibility at the Wikimedia Hackathon 2026

Web accessibility is not merely a technical feature. It is a prerequisite for truly free knowledge. During the recent Wikimedia Hackathon 2026, held in Milan, we came together as a dedicated group hailing from Italy to confront a quiet yet persistent issue: the barriers that prevent visually-impaired individuals from fully engaging with Wikipedia and its sister projects.

Thus, Valcio, Daimona Eaytoy, and Piergiovanna Grossi (WMIT) led the unconference session “Wikipedia for Everyone: Closing the Accessibility Gap”, which served as both a wake-up call and a collaborative workshop. By examining how community-made templates and interface elements often fail our users, we aimed to transition from identifying problems to building sustainable solutions.

This is a short recap for those who missed it.

The Reality of the Digital Barrier

Home page for MediaWiki Accessibility Checker
Home page for MediaWiki Accessibility Checker

The session opened with a candid look at the current state of our interfaces. While MediaWiki provides a robust foundation, years of community-driven customisation have inadvertently introduced many accessibility violations. Key issues discussed included:

  • Missing Alt-Text: Images essential for understanding content often lack descriptions or alternative text which is readable by screen readers, assistive technologies that read out graphic content to visually impaired users.
  • The “HTML Wall”: Many tables and templates lack proper semantic markup, forcing text-to-speech tools to read out raw code rather than structured information.
  • Contrast and Colour: Numerous gadgets and banners still fall short of the WCAG 2.2 AA (a web-accessibility standard) minimum contrast ratios, rendering them invisible to users with colour blindness or low vision.

Measuring Missing Alt-Text

The unconference session also sparked a small follow-up experiment. CristianCantoro set out to measure how widespread the issue of missing alt-text is on Italian Wikipedia and Lombard Wikipedia, combining the Wikipedia HTML dumps provided by Wikimedia Enterprise with the XML dumps published by the Wikimedia Foundation. The initial results confirm the scale of the challenge: more than 90% of images used in Italian and Lombard Wikipedia articles lack alternative text.

This is not an isolated finding. In 2023, a team of researchers from Stanford University and Google Research presented a cross-lingual analysis of image accessibility across 108 Wikipedia language editions finding that, on average, only around 10% of images had alt-text. This research was presented at the 2023 edition of the Wiki Workshop.

These numbers are a reminder that missing alt-text is still an open and large-scale challenge across languages. If we want Wikipedia to be truly open to everyone, we need better tools, workflows, and community practices to help editors add alt-text and meaningful descriptions to images.

From Discussion to Action: The MediaWiki Accessibility Checker

Logo for MediaWiki Accessibility Checker
Logo for MediaWiki Accessibility Checker

To move from awareness to action, one of the session participants — Super nabla from the Indic MediaWiki Developers User Group — built a concrete solution during the hackathon itself. The tool, available on Toolforge, assists editors and developers in meeting accessibility standards: the MediaWiki Accessibility Checker. Try it out: https://accessibility-checker.toolforge.org/

Built on the industry-standard axe-core engine and Playwright, the tool is specifically adapted for the MediaWiki ecosystem. It allows editors and developers to (i) perform deep audits (queryable both from the frontend interface as well as from a dedicated RESTful API) based on WCAG 2.2 AA (and other standards) on any wiki URL, including project pages; (ii) generate professional reports in multiple formats, including PDF and Wikitext for easy sharing on-wiki; (iii) utilise a modern interface designed with the Wikimedia Codex design system, ensuring a seamless experience for contributors.

This tool represents a small yet important step forward in democratising accessibility auditing, allowing gadget authors — even those without formal expertise — to identify and rectify errors before they impact our readers.

A Legacy of “Wikiricci” and Community Care

Daimona Eaytoy with the WikiRiccio
Daimona Eaytoy with the WikiRiccio

The roots of this technical collaboration extend back to 2018 at itWikiCon in Como (Italy), where the “Officina” (the Italian Wikipedia’s technical project) was honoured for its quiet, essential labour, carried out by the smanettoni (hackers) — the tinkerers and wizards who operate behind the scenes to ensure the platform’s gears continue to turn. This community recognition is personified by the Wikiriccio (wiki hedgehog), a physical trophy whose travel history has become something of a legendary saga within the Italian community. Traditionally held in rotation, after years of near-misses, it finally found its way to Daimona Eaytoy during this hackathon, reminding us that accessibility work is also about human connections and shared care.

For us, this light-hearted tradition and award serve as a reminder: behind every accessibility tool or interface fix is a human connection, a shared community-based vision and history, and a commitment to “making the shop run” for the benefit of all users.

Next Steps and Community Involvement

The hackathon session was only the beginning. The outcomes of our session are being synthesised into a formal proposal in the Italian Wikipedia and a Phabricator task to help standardise CSS custom properties and automated linting workflows.

Yet, technology alone cannot solve a cultural challenge. We invite all UI/UX designers, developers, and experienced wiki-editors to join the effort. Whether you are improving the alt text on a high-traffic policy page or helping modernise an old template, your contribution ensures that Wikipedia remains truly accessible, enabling everyone to share in the sum of all knowledge.

A special thanks to the hackathon organisers and all the participants who shared their lived experiences; your insights are what drive these technical improvements forward.

Tech News 2026 – Issue 23

1 June 2026 at 21:17

Latest tech news from the Wikimedia technical community. Please tell other users about these changes. Not all changes will affect you. Translations are available.

Updates for editors

  • The Reader Experience team is conducting an experiment to show the reading lists feature, which is still in development, to logged-out mobile readers to test whether it encourages account creation at a higher rate compared to the watchstar button. The experiment was launched on May 18th on German, Spanish, Italian, Portuguese, Polish, Dutch, Turkish, and Urdu wikis, and it will run for a month.
  • The Wikimedia Apps team released Phase 1 of the redesigned Home Feed to the Android Beta app. The new Home Feed includes a refreshed “Community” tab and a personalized “For You” tab featuring daily updated reading recommendations. The redesign is part of a broader effort to improve content discovery and create more engaging learning experiences in the Wikipedia apps.
  • Recurrent item View all 18 community-submitted tasks that were resolved last week. For example, an issue where images could fail to load for some suggested edits on Special:Homepage, leaving the thumbnail stuck in a loading state, has now been fixed. [1]

Updates for technical contributors

  • Recurrent item Detailed code updates later this week: MediaWiki

Tech news prepared by Tech News writers and posted by bot • Contribute • Translate • Get help • Give feedback • Subscribe or unsubscribe.

AWW Podcast Season 2 Episode #1 Can Wikipedia Evolve With the Digital Age? 

By: AnnComms
1 June 2026 at 07:00

There was a time when Wikipedia was the go-to source for information and one of the most trusted tools for research across the world. From students and journalists to researchers and everyday internet users, millions relied on the platform for quick and accessible knowledge. However, as technology continues to evolve, the way people consume information has also changed.

Today, Wikipedia faces growing competition from emerging technologies such as Artificial Intelligence (AI) tools and social media platforms, which now shape how many people search for and engage with information online. As a result, the platform has experienced a decline in page views over the years, raising important questions about its future relevance and visibility in the digital age.

To address these concerns, about 100 Wikimedian affiliates, volunteers, and external experts gathered in Frankfurt am Main from 30 January to 1 February 2026, for the Wikimedia Futures Lab event organised by the Wikimedia movement. The Futures Lab serves as a space for research, experimentation, and forward-thinking conversations on the future of free knowledge.

At a time when technology is rapidly transforming the internet and information-sharing, the event provided an opportunity for participants to reflect on how Wikipedia can continue to remain relevant, visible, and trusted in an increasingly digital and AI-driven world.

From the attendees

The conversations and ideas shared during the event formed the AWW Voices Podcast episode “Can Wikipedia Evolve with the Digital Age?”. In this episode, host Oluwapelumi Aina joined by Ruby D Brown, Co-Founder of African Wiki Women, Tochi Precious, Language Advocate and Co-Founder of the Igbo User Group, and Olubusola Afolabi, Community Engagement Lead at Free Knowledge Africa. 

Screenshot of AWW Voices Podcast host and guests.

Having attended the Wikimedia Futures Lab event, the guests shared their experiences, reflections, and key takeaways from the discussions held in Frankfurt. 

“The world around us is changing really fast. When you think about how people trust information online, AI-generated media, new laws, and shifting technologies, it becomes important to understand how these trends affect us as the Wikimedia community,” says Tochi.

Wikipedia vs Digital Age

Despite technological advancement, Wikipedia, once regarded as one of the most trusted digital information platforms, has seen a decline in page views since 2016 as more people turn to AI tools for information. However, it is important to recognise that many AI systems are trained using content from platforms like Wikipedia.

“For example, when you search for something on Google, the AI overview provides a summary alongside references. Very few people actually click on the Wikipedia link for the longer version. This shows that people are still consuming Wikipedia content, but AI tools now act as middlemen,” explains Olubusola.

According to her, this shift means Wikipedia can no longer rely solely on users visiting the platform directly. Instead, it must adapt to changing online habits and find ways to bring information closer to the spaces where audiences already spend their time.

She adds that Wikipedia must adapt by meeting audiences where they already are, bringing information directly to the platforms people use instead of expecting them to always visit the main website.

The solution

The rise of AI and social media has also changed how people consume information. Many users now prefer short-form content over long-form reading because of shrinking attention spans. Since Wikipedia is traditionally a long-form platform, there is growing pressure for it to evolve alongside these changing habits.

For many younger internet users, information is no longer consumed through lengthy articles alone. Videos, creators, podcasts, and short-form explainers are increasingly becoming the preferred way to learn and engage online.

“People are moving away from institution-based information and increasingly relying on personalities. They want direct interaction, and video content makes information easier to consume. As Wikimedia, we need to pay attention to these shifts so we can meet people where they are,” says Ruby.

The Dilemma

Wikimedia exists because of the volunteers who edit and write the content on the platform. While keeping up with technological change is necessary, the movement also faces the challenge of ensuring that technology does not overshadow the human element that has always been at the centre of Wikimedia projects.

As conversations around AI continue to grow, many community members believe the focus should remain on supporting contributors rather than replacing them.

Last year, the Wikimedian community launched its AI Strategy, which clearly showed that AI should not replace the human writers and editors but rather support their work.

When the Home Page Gets Boring: How My Colleagues and I Revitalised Thai Wikipedia

31 May 2026 at 18:00

After a few years away from Thai Wikipedia, I returned to find that the Main Page had become stagnant. It lacked the dynamic energy a landing page needs. So, my colleagues and I decided to revitalise it—and here is exactly how we did it.

Thai Wikipedia's Home Page, as of 26 May 2026, only the website's logo, search box, page name. welcome message, featured sections and broad categories links included.
Thai Wikipedia’s Home Page, as of 26 May 2026

Before diving into the details, let me explain the structure of Thai Wikipedia’s Home Page. It was heavily inspired by the original English edition‘s layout, featuring four core content sections:

  • This Month’s Featured Articles (TMFA): An excerpt of a well-written article (Thai Wikipedia lacks the volume to change this daily like the English site).
  • Did You Know (DYK): Interesting facts pulled from recently expanded or created articles.
  • In The News (ITN): Recent global (and occasionally space-related) events.
  • On This Day (OTD): A look back at historical events on the current date.

When I returned to active editing in mid-2024, I realised these sections were frozen in time. Sometimes, content remained identical for days. After a thorough review, I found the issues were threefold: stagnant content, unpredictable update schedules (except for the strictly automated OTD), and complex, opaque backend procedures for publishing content to the Main Page.

To build a sustainable solution, we had to attack the problem from two angles: community contribution and technical infrastructure.

On the contribution side, we introduced clear, easy-to-follow Standard Operating Procedures (SOPs) to ensure nominators and reviewers wouldn’t feel overwhelmed. We also lifted several legacy constraints that were discouraging newbie and intermediate editors.

On the nerdy side, we introduced a “Nested Transclude Template System” to make pulling content to the main page seamless. No more messy, bespoke coding required. All nominations can now be tracked and recalled without digging through a chaotic page history.

For the less tech-savvy, here is how simple it is now: You no longer need to deal with any messy, complicated coding. As shown in the diagram, everything is built like a set of nesting dolls:

Diagram illustrating a nested template system for Wikipedia. Content like hooks and excerpts are grouped inside date-based templates, which are automatically pulled into the main DYK and TMRA templates.
A Diagram to demonstrate a nested template system for Wikipedia. Content like hooks and excerpts are grouped inside date-based templates, which are automatically pulled into the main DYK and TMRA templates.
  • Write your content: You just write your proposal or excerpt in a standard form.
  • Name it with the date: You save it inside a specific date format (like YYYY-MM-DD).
  • The system does the rest: When that day arrives, the Main Page template automatically fetches the correct date’s content and puts it live—completely on its own!

This means no one has to lift a finger to update it manually, and we can track past nominations without digging through a chaotic page history.

Did You Know it’s now easier than ever to nominate your articles?

The first backlog I tackled was the DYK section. There, I crossed paths with Taweethaも, a renowned Thai Wikipedian. That chance encounter inspired a complete revolution of our process. We teamed up to clear backlogs that had been sitting untouched for over six months. Together, we drafted new SOPs and built a backend system to support them—queuing content chronologically by nomination date, enforcing character limits, and scheduling release dates.

Once the system stabilised, we launched a content contest to diversify the topics and test our new workflow under pressure. The campaign was a massive success: 16 contributors created or improved over 90 articles. Crucially, three of those contributors remain highly active “DYK editors” today.

We also noticed that while some nominators were incredibly prolific, they rarely helped review others’ work. To keep the backlog manageable, we implemented a Quid Pro Quo (QPQ) policy, requiring nominators to review a peer’s submission to qualify their own.

Opening the Gates: Allowing Good Articles onto the Main Page

With DYK running smoothly, we turned our attention to TMFA. This section had suffered from a decade-long drought of new Featured Articles (FAs) to showcase. Beyond adapting our new DYK SOPs, we made a major policy shift: we lifted the strict FA constraint and allowed Good Articles (GAs) to be featured. To reflect this, we renamed the section from This Month’s Featured Article to Recommended Articles.

Whilst long-form, high-quality writing requires significantly more energy from contributors—meaning it wasn’t as explosive as the DYK campaign—the initiative still successfully brought 7 brand-new, high-quality articles to the front page from 7 different writers.

A new solution brings a new quirk

Excerpt of Thai Wikipedia's Home Page on 4 June 2025, but it displayed OTD of 31 May.
An excerpt of Thai Wikipedia’s Home Page on 4 June 2025 showing OTD content from 31 May due to caching issues.

Every new system has its bugs. Just a day into the DYK campaign, a participant noticed that logged-out readers were seeing stale, outdated main page content, while logged-in users saw the updates perfectly.

We spent days hunting for a fix. Thankfully, User:Chlod—a perennial savior of Wikipedia infrastructure—pointed out that the server cache just needed to be manually “purged” (which simply means appending ?action=purge to the URL string).

To automate this, I sat down for some classic “vibe coding” and wrote a Python script. Hosted on Toolforge (Wikimedia’s dedicated server for customised scripts within the Wikimedia Movement) and linked to my bot account, it now runs via a cron job twice a day to keep the page fresh. I also added a secondary feature to the script: it automatically archives the Main Page to the Internet Archive‘s Wayback Machine daily.

For those unfamiliar with the tech jargon, here is the simple version: I asked the AI chatbot, Google Gemini, to help me write a program in the Python language. After testing it repeatedly until I was sure it worked, I uploaded the code to Toolforge—which is essentially a free, 24/7 computer server available to Wikipedia volunteers. I set the server to run my code twice a day to automatically fix the glitch and keep the Main Page fresh. As a bonus, I also programmed it to save a daily copy of the Main Page to the Wayback Machine (a digital archive of the internet) so we always have a historical record.

I’ve published my source code in GitHub if you’re looking for: https://github.com/sarawutkhs/wthpurge

What about the other two sections?

You might be wondering why I haven’t mentioned ITN or OTD. To be completely honest, I tried to implement similar reforms for OTD, but couldn’t find anyone in the community available to jump in. If you have ideas on how we can spark interest and bring that same magic to the remaining sections, please drop a comment!

Acknowledgements

This transformation wouldn’t have been possible without an incredible support system. Beyond those already mentioned, I want to thank the original architects of the Main Page structure, as well as every single campaign participant who dedicated time to improving Thai Wikipedia. Finally, my deepest respect goes to Taweethaも, whose guidance both on- and off-wiki was invaluable.

Declaration: This case study was previously presented at the ESEAP Conference 2026 and the October 2025 ESEAP Community Call. The initial phase of this project was also published on the ESEAP’s Substack.

A surprising IC in a LED light chain.

By: cpldcpu
25 November 2024 at 19:23

LED-based festive decorations are a fascinating subject for exploration of ingenuity in low-cost electronics. New products appear every year and often very surprising technology approaches are used to achieve some differentiation while adding minimal cost.

This year, there wasn’t any fancy new controller, but I was surprised how much the cost of simple light strings was reduced. The LED string above includes a small box with batteries and came in a set of ten for less than $2 shipped, so <$0.20 each. While I may have benefitted from promotional pricing, it is also clear that quite some work went into making the product cheap.

The string is constructed in the same way as one I had analyzed earlier: it uses phosphor-converted blue LEDs that are soldered to two insulated wires and covered with an epoxy blob. In contrast to the earlier device, they seem to have switched from copper wire to cheaper steel wires.

The interesting part is in the control box. It comes with three button cells, a small PCB, and a tactile button that turns the string on and cycles through different modes of flashing and and constant light.

Curiously, there is nothing on the PCB except the button and a device that looks like an LED. Also, note how some “redundant” joints have simply been left unsoldered.

Closer inspection reveals that the “LED” is actually a very small integrated circuit packaged in an LED package. The four pins are connected to the push button, the cathode of the LED string, and the power supply pins. I didn’t measure the die size exactly, but I estimate that it is smaller than 0.3×0.2 mm² = ~0.1 mm².

What is the purpose of packaging an IC in an LED package? Most likely, the company that made the light string is also packaging their own LEDs, and they saved costs by also packaging the IC themselves—in a package type they had available.

I characterized the current-voltage behavior of IC supply pins with the LED string connected. The LED string started to emit light at around 2.7V, which is consistent with the forward voltage of blue LEDs. The current increased proportionally to the voltage, which suggests that there is no current limit or constant current sink in the IC – it’s simply a switch with some series resistance.

Left: LED string in “constantly on” mode. Right: Flashing

Using an oscilloscope, I found that the string is modulated with an on-off ratio of 3:1 at a frequency if ~1.2 kHz. The image above shows the voltage at the cathode, the anode is connected to the positive supply. This is most likely to limit the current.

All in all, it is rather surprising to see an ASIC being used when it barely does more than flashing the LED string. It would have been nice to see a constant current source to stabilize the light levels over the lifetime of the battery and maybe more interesting light effects. But I guess that would have increased the cost of the ASIC too much and then using an ultra-low cost microcontroller may have been cheaper. This almost calls for a transplant of a MCU into this device…

Neural Networks (MNIST inference) on the “3-cent” Microcontroller

By: cpldcpu
2 May 2024 at 23:59

Bouyed by the surprisingly good performance of neural networks with quantization aware training on the CH32V003, I wondered how far this can be pushed. How much can we compress a neural network while still achieving good test accuracy on the MNIST dataset? When it comes to absolutely low-end microcontrollers, there is hardly a more compelling target than the Padauk 8-bit microcontrollers. These are microcontrollers optimized for the simplest and lowest cost applications there are. The smallest device of the portfolio, the PMS150C, sports 1024 13-bit word one-time-programmable memory and 64 bytes of ram, more than an order of magnitude smaller than the CH32V003. In addition, it has a proprieteray accumulator based 8-bit architecture, as opposed to a much more powerful RISC-V instruction set.

Is it possible to implement an MNIST inference engine, which can classify handwritten numbers, also on a PMS150C?

On the CH32V003 I used MNIST samples that were downscaled from 28×28 to 16×16, so that every sample take 256 bytes of storage. This is quite acceptable if there is 16kb of flash available, but with only 1 kword of rom, this is too much. Therefore I started with downscaling the dataset to 8×8 pixels.

The image above shows a few samples from the dataset at both resolutions. At 16×16 it is still easy to discriminate different numbers. At 8×8 it is still possible to guess most numbers, but a lot of information is lost.

Suprisingly, it is still possible to train a machine learning model to recognize even these very low resolution numbers with impressive accuracy. It’s important to remember that the test dataset contains 10000 images that the model does not see during training. The only way for a very small model to recognize these images accurate is to identify common patterns, the model capacity is too limited to “remember” complete digits. I trained a number of different network combinations to understand the trade-off between network memory footprint and achievable accuracy.

Parameter Exploration

The plot above shows the result of my hyperparameter exploration experiments, comparing models with different configurations of weights and quantization levels from 1 to 4 bit for input images of 8×8 and 16×16. The smallest models had to be trained without data augmentation, as they would not converge otherwise.

Again, there is a clear relationship between test accuracy and the memory footprint of the network. Increasing the memory footprint improves accuracy up to a certain point. For 16×16, around 99% accuracy can be achieved at the upper end, while around 98.5% is achieved for 8×8 test samples. This is still quite impressive, considering the significant loss of information for 8×8.

For small models, 8×8 achieves better accuracy than 16×16. The reason for this is that the size of the first layer dominates in small models, and this size is reduced by a factor of 4 for 8×8 inputs.

Surprisingly, it is possible to achieve over 90% test accuracy even on models as small as half a kilobyte. This means that it would fit into the code memory of the microcontroller! Now that the general feasibility has been established, I needed to tweak things further to accommodate the limitations of the MCU.

Training the Target Model

Since the RAM is limited to 64 bytes, the model structure had to use a minimum number of latent parameters during inference. I found that it was possible to use layers as narrow as 16. This reduces the buffer size during inference to only 32 bytes, 16 bytes each for one input buffer and one output buffer, leaving 32 bytes for other variables. The 8×8 input pattern is directly read from the ROM.

Furthermore, I used 2-bit weights with irregular spacing of (-2, -1, 1, 2) to allow for a simplified implementation of the inference code. I also skipped layer normalization and instead used a constant shift to rescale activations. These changes slightly reduced accuracy. The resulting model structure is shown below.

All things considered, I ended up with a model with 90.07% accuracy and a total of 3392 bits (0.414 kilobytes) in 1696 weights, as shown in the log below. The panel on the right displays the first layer weights of the trained model, which directly mask features in the test images. In contrast to the higher accuracy models, each channel seems to combine many features at once, and no discernible patterns can be seen.

Implementation on the Microntroller

In the first iteration, I used a slightly larger variant of the Padauk Microcontrollers, the PFS154. This device has twice the ROM and RAM and can be reflashed, which tremendously simplifies software development. The C versions of the inference code, including the debug output, worked almost out of the box. Below, you can see the predictions and labels, including the last layer output.

Squeezing everything down to fit into the smaller PMS150C was a different matter. One major issue when programming these devices in C is that every function call consumes RAM for the return stack and function parameters. This is unavoidable because the architecture has only a single register (the accumulator), so all other operations must occur in RAM.

To solve this, I flattened the inference code and implemented the inner loop in assembly to optimize variable usage. The inner loop for memory-to-memory inference of one layer is shown below. The two-bit weight is multiplied with a four-bit activation in the accumulator and then added to a 16-bit register. The multiplication requires only four instructions (t0sn, sl,t0sn,neg), thanks to the powerful bit manipulation instructions of the architecture. The sign-extending addition (add, addc, sl, subc) also consists of four instructions, demonstrating the limitations of 8-bit architectures.

void fc_innerloop_mem(uint8_t loops) {

    sum = 0;
    do  {
       weightChunk = *weightidx++;
__asm   
    idxm  a, _activations_idx
	inc	_activations_idx+0

    t0sn _weightChunk, #6
    sl     a            ;    if (weightChunk & 0x40) in = in+in;
    t0sn _weightChunk, #7
    neg    a           ;     if (weightChunk & 0x80) in =-in;                    

    add    _sum+0,a
    addc   _sum+1
    sl     a 
    subc   _sum+1  

  ... 3x more ...

__endasm;
    } while (--loops);

    int8_t sum8 = ((uint16_t)sum)>>3; // Normalization
    sum8 = sum8 < 0 ? 0 : sum8; // ReLU
    *output++ = sum8;
}

In the end, I managed to fit the entire inference code into 1 kilowords of memory and reduced sram usage to 59 bytes, as seen below. (Note that the output from SDCC is assuming 2 bytes per instruction word, while it is only 13 bits).

Success! Unfortunately, there was no rom space left for the soft UART to output debug information. However, based on the verificaiton on PFS154, I trust that the code works, and since I don’t have any specific application in mind, I left it at that stage.

Summary

It is indeed possible to implement MNIST inference with good accuracy using one of the cheapest and simplest microcontrollers on the market. A lot of memory footprint and processing overhead is usually spent on implementing flexible inference engines, that can accomodate a wide range of operators and model structures. Cutting this overhead away and reducing the functionality to its core allows for astonishing simplification at this very low end.

This hack demonstrates that there truly is no fundamental lower limit to applying machine learning and edge inference. However, the feasibility of implementing useful applications at this level is somewhat doubtful.

You can find the project repository here.

Neural Networks (MNIST inference) on the “3-cent” Microcontroller

By: cpldcpu
2 May 2024 at 23:59

Bouyed by the surprisingly good performance of neural networks with quantization aware training on the CH32V003, I wondered how far this can be pushed. How much can we compress a neural network while still achieving good test accuracy on the MNIST dataset? When it comes to absolutely low-end microcontrollers, there is hardly a more compelling target than the Padauk 8-bit microcontrollers. These are microcontrollers optimized for the simplest and lowest cost applications there are. The smallest device of the portfolio, the PMS150C, sports 1024 13-bit word one-time-programmable memory and 64 bytes of ram, more than an order of magnitude smaller than the CH32V003. In addition, it has a proprieteray accumulator based 8-bit architecture, as opposed to a much more powerful RISC-V instruction set.

Is it possible to implement an MNIST inference engine, which can classify handwritten numbers, also on a PMS150C?

On the CH32V003 I used MNIST samples that were downscaled from 28×28 to 16×16, so that every sample take 256 bytes of storage. This is quite acceptable if there is 16kb of flash available, but with only 1 kword of rom, this is too much. Therefore I started with downscaling the dataset to 8×8 pixels.

The image above shows a few samples from the dataset at both resolutions. At 16×16 it is still easy to discriminate different numbers. At 8×8 it is still possible to guess most numbers, but a lot of information is lost.

Suprisingly, it is still possible to train a machine learning model to recognize even these very low resolution numbers with impressive accuracy. It’s important to remember that the test dataset contains 10000 images that the model does not see during training. The only way for a very small model to recognize these images accurate is to identify common patterns, the model capacity is too limited to “remember” complete digits. I trained a number of different network combinations to understand the trade-off between network memory footprint and achievable accuracy.

Parameter Exploration

The plot above shows the result of my hyperparameter exploration experiments, comparing models with different configurations of weights and quantization levels from 1 to 4 bit for input images of 8×8 and 16×16. The smallest models had to be trained without data augmentation, as they would not converge otherwise.

Again, there is a clear relationship between test accuracy and the memory footprint of the network. Increasing the memory footprint improves accuracy up to a certain point. For 16×16, around 99% accuracy can be achieved at the upper end, while around 98.5% is achieved for 8×8 test samples. This is still quite impressive, considering the significant loss of information for 8×8.

For small models, 8×8 achieves better accuracy than 16×16. The reason for this is that the size of the first layer dominates in small models, and this size is reduced by a factor of 4 for 8×8 inputs.

Surprisingly, it is possible to achieve over 90% test accuracy even on models as small as half a kilobyte. This means that it would fit into the code memory of the microcontroller! Now that the general feasibility has been established, I needed to tweak things further to accommodate the limitations of the MCU.

Training the Target Model

Since the RAM is limited to 64 bytes, the model structure had to use a minimum number of latent parameters during inference. I found that it was possible to use layers as narrow as 16. This reduces the buffer size during inference to only 32 bytes, 16 bytes each for one input buffer and one output buffer, leaving 32 bytes for other variables. The 8×8 input pattern is directly read from the ROM.

Furthermore, I used 2-bit weights with irregular spacing of (-2, -1, 1, 2) to allow for a simplified implementation of the inference code. I also skipped layer normalization and instead used a constant shift to rescale activations. These changes slightly reduced accuracy. The resulting model structure is shown below.

All things considered, I ended up with a model with 90.07% accuracy and a total of 3392 bits (0.414 kilobytes) in 1696 weights, as shown in the log below. The panel on the right displays the first layer weights of the trained model, which directly mask features in the test images. In contrast to the higher accuracy models, each channel seems to combine many features at once, and no discernible patterns can be seen.

Implementation on the Microntroller

In the first iteration, I used a slightly larger variant of the Padauk Microcontrollers, the PFS154. This device has twice the ROM and RAM and can be reflashed, which tremendously simplifies software development. The C versions of the inference code, including the debug output, worked almost out of the box. Below, you can see the predictions and labels, including the last layer output.

Squeezing everything down to fit into the smaller PMS150C was a different matter. One major issue when programming these devices in C is that every function call consumes RAM for the return stack and function parameters. This is unavoidable because the architecture has only a single register (the accumulator), so all other operations must occur in RAM.

To solve this, I flattened the inference code and implemented the inner loop in assembly to optimize variable usage. The inner loop for memory-to-memory inference of one layer is shown below. The two-bit weight is multiplied with a four-bit activation in the accumulator and then added to a 16-bit register. The multiplication requires only four instructions (t0sn, sl,t0sn,neg), thanks to the powerful bit manipulation instructions of the architecture. The sign-extending addition (add, addc, sl, subc) also consists of four instructions, demonstrating the limitations of 8-bit architectures.

void fc_innerloop_mem(uint8_t loops) {

    sum = 0;
    do  {
       weightChunk = *weightidx++;
__asm   
    idxm  a, _activations_idx
	inc	_activations_idx+0

    t0sn _weightChunk, #6
    sl     a            ;    if (weightChunk & 0x40) in = in+in;
    t0sn _weightChunk, #7
    neg    a           ;     if (weightChunk & 0x80) in =-in;                    

    add    _sum+0,a
    addc   _sum+1
    sl     a 
    subc   _sum+1  

  ... 3x more ...

__endasm;
    } while (--loops);

    int8_t sum8 = ((uint16_t)sum)>>3; // Normalization
    sum8 = sum8 < 0 ? 0 : sum8; // ReLU
    *output++ = sum8;
}

In the end, I managed to fit the entire inference code into 1 kilowords of memory and reduced sram usage to 59 bytes, as seen below. (Note that the output from SDCC is assuming 2 bytes per instruction word, while it is only 13 bits).

Success! Unfortunately, there was no rom space left for the soft UART to output debug information. However, based on the verificaiton on PFS154, I trust that the code works, and since I don’t have any specific application in mind, I left it at that stage.

Summary

It is indeed possible to implement MNIST inference with good accuracy using one of the cheapest and simplest microcontrollers on the market. A lot of memory footprint and processing overhead is usually spent on implementing flexible inference engines, that can accomodate a wide range of operators and model structures. Cutting this overhead away and reducing the functionality to its core allows for astonishing simplification at this very low end.

This hack demonstrates that there truly is no fundamental lower limit to applying machine learning and edge inference. However, the feasibility of implementing useful applications at this level is somewhat doubtful.

You can find the project repository here.

Decapsulating the CH32V203 Reveals a Separate Flash Die

By: cpldcpu
1 May 2024 at 11:02

The CH32V203 is a 32bit RISC-V microcontroller. In the produt portfolio of WCH it is the next step up from the CH32V003, sporting a much higher clock rate of 144 MHz and a more powerful RISC-V core with RV32IMAC instruction set architecture. The CH32V203 is also extremely affordable, starting at around 0.40 USD (>100 bracket), depending on configuration.

An interesting remark on twitter piqued my interest: Supposedly the listed flash memory size only refers to a fraction that can be accessed with zero waitstate, while the total flash size is even 224kb. The datasheet indeed has a footnote claiming the same. In addition, the RB variant offers the option to reconfigure between RAM and flash, which is rather odd, considering that writing to flash is usually much slower than to RAM.

Then the 224kb number is mentioned in the memory map. Besides the code flash, there is also a 28Kb boot section and additional configurable space. 224 kbyte +28 kbyte+4=256kbyte, which suggests that the total available flash is 256 kbyte and is remapped to different locations of the memory.

All of these are red flags for an architecture where a separate NOR flash die is used to store the code and the main CPU core has a small SRAM that is used as a cache. This configuration was pioneered by Gigadevice and is also famously used by the ESP32 and RP2040 more recently, although that latter two use an external NOR flash device.

Flash memory is quite different from normal CMOS devices as it requires a special gate stack, isolation and much higher voltages. Therefore, integrating flash memory into a CMOS logic die usually requires extra process steps. The added complexity increases when going to smaller technologies nodes. Separating both dies offers the option of using a high density logic process (for example 45 nm) and pairing it with a low-cost off-the-shell NOR flash die.

Decapsulation and Die Images

To confirm my suspicions I decapsulated a CH32V203C8T6 sample, shown above. I heated the package to drive out the resin and then carefully broke the, now brittle, package apart. Already after removing the lead frame, we can cleary see that it contains two dies.

The small die is around 0.5mm² in area. I wasn’t able to completely removed the remaining filler, but we can see that it is an IC with a smaller number of pads, fitting to a serial flash die.

The microcontroller die came out really well. Unfortunately, the photos below are severely limited by my low-cost USB microscope. I hope Zeptobars or others will come up with nicer images at some point.

The die size of ~1.8 mm² is surprisingly small. In fact it is even smaller than the die of the CH32V003 with a die size of ~2.0 mm² according to Zeptobars die shot. Apart from the fact that the flash was moved off-chip, most likely also a much smaller CMOS technology node was used for the CH32V203 than for the V003.

Summary

It was quite surprising to find a two-die configuration in such a low-cost device. But obviously, it explains the oddities in the device specification, and it also explains why 144 MHz core clock is possible in this device without wait-states.

What are the repercussions?

Amazingly, it seems that, instead of only 32kb of flash, as listed for the smallest device, a total of 224kb can be used for code and data storage. The datasheet mentions a special “flash enhanced read mode” that can apparently be used to execute code from the extended flash space. It’s not entirely clear what the impact on speed is, though, but that’s certainly an area for exploration.

I also expect this MCU to be highly overclockable, similar to the RP2040.

What are you really doing when you fill in an hCaptcha

8 January 2022 at 01:00

hCaptcha is a reCAPTCHA clone that has been growing in popularity over 2020 and 2021, in particular due to Cloudflare’s conversion of their nag screens from Google’s reCAPTCHA to hCaptcha. Although hCaptcha advertises itself as being a privacy-conscious alternative to reCAPTCHA, there’s also an incentive for websites to switch over: hCaptcha will pay websites each time one of their users completes a hCaptcha challenge.

Now the question is: how does you completing a captcha earn anyone money? Of course, hCaptcha is a VC-funded business, so it can afford to burn money in the pursuit of market share; nonetheless there needs to be a plausible business model there, and it’s not obvious at first sight.

If you read the hCaptcha website, they suggest that AI startups will pay them to label their images for them. 1 Labelling images is a labour-intensive task and required for some current-generation machine learning approaches. AI startups are well-funded and have money to spend on labelling, so this sounds like a reasonable case of selling shovels during a gold rush. But the output from solving CAPTCHAs isn’t obviously isomorphic to the type of labelling required for machine learning, which is often quite specific and requires a very low error rate.

Complex CAPTCHA challenges are not possible, as web users turn out to be drunk, blind, 3 years old, or just randomly clicking buttons to get this infernal thing to go away. Accordingly, hCaptcha challenges are simple: select the images that match a simple 1-3 word prompt from a 3x3 grid. This is fortunately easy for most real people. 2 3

The most common prompts seem to be selecting buses, trucks, boats or trains out of the grid.4 The market demand for this sort of simple labelling must be rather limited, even if challenges have to be repeated many times and cross-checked to get an acceptable error rate.

So far, a little inscrutable but all seems sensible enough. But then it all gets interesting when you actually take a look at the images in a little more detail:

hCaptcha example

Starting from the top left and going right, we have:

  • A boat that appears to have been painted by Dalí, with a mast drooping like a wet noodle.
  • A plane with tricycle landing gear, except it’s got two sets of wheels at the front and one at the back. That’s not normal!
  • A normal looking plane with some odd-looking clouds above.
  • A bus with an axle in front of the door, and another behind it, and another at the back. Hmm
  • A boat in a marina made of splodges.
  • A normal-looking boat on a normal-looking sea, except - look at that horizon! How did that happen.
  • A single-decker london bus with a ghost of it’s double-decker cousin above. And a giant moth perched on it at the back.
  • Another ghostly upper deck on a regional bus.
  • A sailing boat with some oddly stylised “alien” writing on the sail.

These images are obviously AI-generated. They have all the hallmarks of GAN output, with typical artifacts and oddities. Have some more and see if you can spot the same things in these other challenges - it’s not hard at all, is it!

The question then is why? Why would hCaptcha be generating these challenges - aren’t they supposed to be labelling real life, not some AI mirages? You know the labels before you generate them, what’s the point in using humans to re-label them again… And why are the results so bad - these are definitely not state of the art!

The only explanation that makes sense is that hCaptcha is not really doing this whole AI-labelling business at all, or if they are it’s only in very limited fashion. Most of the time they’re just using a GAN to generate images that defeat the bots’ image recognition AI. And the GAN isn’t trained to optimise human recognition, rather to confound the bots in an arms race, leading to the bad image quality.

If you have any better ideas I’d be glad to hear them because this whole thing doesn’t really make much sense.

Footnotes:

  1. If you look closer, they have an article that purports to explain the “technical architecture of hCaptcha” which is a supreme example of buzzword-stuffing blockchain-washed nothing. There is less than zero need for a blockchain to track customer requests, much less the public Ethereum blockchain, but it’s the buzzword of the month so it must go in. 

  2. Most real users, that is. There are some users for whom the challenge is actually too hard, or who’ve been blackholed and are interpreting bad IP reputation as poor skill. But the ones who fall down most often are those who try too hard and analyse the prompt and challenge in too much detail. The real way to solve these image challenges is to answer what you think other people will answer, rather than the correct answer. And don’t take too long either, just a quick glance is all your competition are giving! Anecdotally, this isn’t too common with hCaptcha, but reCAPTCHA challenges are extremely prone to this failure if you think too hard. 

  3. Unfortunately this is also quite easy for bots, somewhat subverting the point of a CAPTCHA, so that’s how browser fingerprinting and IP reputation creep in to get reasonable enough results. 

  4. These prompts are so common that a front-page post on Hacker News consisted of this observation (and prompted me to write up my thoughts on the topic from the past few months). 

Searching for Nothing, Finding a Surprise

23 December 2021 at 21:30

Following on from my post yesterday about an edge case in YouTube, I thought I’d write about a class of edge cases perhaps even more strange that I’ve been exploring recently:

Search engines are a fact of daily life for most of the population nowadays. Google (sub your preferred provider) is an extension of the brain, imagined as giving you access to the sum of the world’s information at the click of a button. But a search engine isn’t just a Ctrl-F for the internet with a nice interface and ads; rather it’s a tremendously complicated system with lots of features and interactions between those features. And all you need to explore the system yourself is some well-tuned search queries.

I recently had an epiphany: search engines are designed to find you results for something and that’s a job they perform well. But there’s nothing stopping you from searching for nothing! And the search engines will still give you results!

And what results they are - have a go on the links below:

An empty query on DDG: https://duckduckgo.com/?q=+””
A different empty query on DDG: https://duckduckgo.com/?q=(“”)
An empty query on Google: https://www.google.com/search?q=(“”)
An empty query on Google News: https://www.google.com/search?q=”“&tbm=nws

And have you ever thought about doing an anything but search? Normally you can add negations to the end of your search term to remove unwanted results, but there’s nothing stopping you from having a search term consisting entirely of negations!

Here’s one on DDG: https://duckduckgo.com/?q=-“an entirely negated query”
On Bing: https://www.bing.com/search?q=-“an entirely negated query”
And on Google Books: https://www.google.com/search?q=-“nothing to see here”&tbm=bks

Commentary

Google appears to have some half-effective filtering for these empty search queries so you’ll mostly get the same two YouTube videos as a result - is this an Easter egg? Although Google News and Books don’t have any filter, and you do get some odd results there!

DuckDuckGo doesn’t appear to have any filtering at all, although it’s obvious just how much DDG relies on Bing’s whitelabel product for its results by looking at how similar the two are.

If you can think of a deeper reason for these results, please do leave a comment and lets try and explain some of the mystery away.

❌
❌