TL;DR: In the last year, the Wikimedia Foundation has fired several union organizers, including those that worked on the Community Tech team - a team dedicated to building features for the volunteer community that edits Wikipedia.

As the Wiki Workers Union tries to get the Wikimedia Foundation to recognize their union, it is worth remembering that this is not the first time that the Foundation has worked against the community.

Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus - a body of work that the Foundation does not own.

Movement volunteer communities contributed to Wikipedia under copyleft licenses - licenses that work to ensure that the work remains free (as in speech). The big tech AI companies do not license derivative works under copyleft licenses and often do not even attribute where the works came from.

This means that volunteers are working for big tech for free, and the Wikimedia Foundation is selling privileged access to that free labor.

It is against that backdrop that the current unionization struggle unfolds.

  • Aatube@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    4
    ·
    2 days ago

    There definitely needs to be more awareness of the union issues, but the Enterprise part of the post body is quite misguided.

    Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus

    This omits how scrapers and crawlers would have been getting the corpus for free without that, profiteering and clogging up the tubes that bring you Wikipedia and tons of amazing applications that rely on it. API access is extraordinarily free and public! I’d rather Google pay for it so they don’t slow down literally everybody else. I for one am glad I don’t need to Anubis every 12h to access something as essential as Wikipedia. The community broadly supports Wikimedia Enterprise too.

    • yoasif@fedia.ioOP
      link
      fedilink
      arrow-up
      1
      ·
      2 days ago

      This omits how scrapers and crawlers would have been getting the corpus for free without that, profiteering and clogging up the tubes that bring you Wikipedia and tons of amazing applications that rely on it.

      I didn’t know that I had to preemptively defend the Foundation.

      The reason I don’t think that that is particularly relevant is because these companies are violating the license that Wikimedia projects are distributed under. WMF has $296M in reserves. They couldn’t send a Cease and Desist for the server load?

      • Aatube@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 days ago

        They couldn’t send a Cease and Desist for the server load?

        1. Send one to whom? A WHOIS of the Brazilian IP that turns out to be a residential proxy? Anonymous scraping for LLMs is a problem all over the Internet without a solution (save Anubis). There’s a reason all the big companies have went for Cloudflare instead of any lawsuits, which you’d hear in the news.
        2. Google has always been correctly attributing Wikipedia content in their info cards before the rise of LLMs and Wikimedia Enterprise has always been made for that purpose. With how much they impact the server load while also being a big driver a traffic, Google clearly needed to be moved off the public infrastructure but it would also be unreasonable to shut them off and starve the site off a big source of contributors. Paying for better service beyond reasonable, individual, volunteer use has been common practice since Red Hat and commercial support, and now Tidelift. See also Grafana and the OpenStreetMap ecosystem.
        3. They’re not even sure if they have the standing; they’ve looked at that question:

        Overall, it is more likely than not if current precedent holds that training systems on copyrighted data will be covered by fair use in the United States, but there is significant uncertainty at time of writing.

        https://meta.wikimedia.org/wiki/Wikilegal/Copyright_Analysis_of_ChatGPT

        I didn’t know that I had to preemptively defend the Foundation.

        well you shouldn’t either. I am simply representing the perspective of contributors as a signatory myself, and as a developer who makes API calls to Wikipedia. “Our content is always free to use, but our infrastructure is not” sums it up nicely and Wikimedia Enterprise keeps the free infrastructure good.

        • yoasif@fedia.ioOP
          link
          fedilink
          arrow-up
          1
          ·
          2 days ago

          Send one to whom? A WHOIS of the Brazilian IP that turns out to be a residential proxy? Anonymous scraping for LLMs is a problem all over the Internet without a solution (save Anubis). There’s a reason all the big companies have went for Cloudflare instead of any lawsuits, which you’d hear in the news.

          I know individuals that have been unmasked in torrent swarms and have had their ISP cancel their service due to that. The idea that Wikimedia’s hands are powerless to send a Cease and Desist to ISPs to warn and ban their customers for scraping is hilarious.

          Again, they have $296M in reserves. WMF can send a letter.

          They’re not even sure if they have the standing; they’ve looked at that question

          But that isn’t what you linked to says; the word standing doesn’t appear in the text, nor do they seem to explore that. Thanks for the reference, but it doesn’t actually support your argument.

          I am simply representing the perspective of contributors as a signatory myself, and as a developer who makes API calls to Wikipedia. “Our content is always free to use, but our infrastructure is not” sums it up nicely and Wikimedia Enterprise keeps the free infrastructure good.

          As I responded to another commenter:

          I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

          Enabling non-violating use cases don’t erase the violating ones.

          • Aatube@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            1
            ·
            2 days ago

            individuals that have been unmasked in torrent swarms

            WMF isn’t Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months’ rain fund in case something massive happens to current revenue.

            But that isn’t what you linked to says; the word standing doesn’t appear in the text

            “Standing” means you have been harmed by an illegal act. You need standing to cease and desist or sue. To have legal standing for a letter, what you’re asking to cease and desist needs to be illegal. What I linked to is that WMF counsel believes LLMs likely do not violate the content licenses of their training data. Thus, there is no legal injury and no standing.

            Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally

            see what I quoted. If we go by ethics (which I do prefer) instead of the law, we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists—in fact, they have been happening despite Enterprise (although at a reduced rate) to the point where rate limits have been enacted for the first time in history—so they might as well make the exploiters pay for it to help maintain the commons’ infrastructure.

            • yoasif@fedia.ioOP
              link
              fedilink
              arrow-up
              1
              ·
              2 days ago

              To have legal standing for a letter, what you’re asking to cease and desist needs to be illegal.

              FWIW, that isn’t true - I didn’t post the letter, but I got a C&D from SoFi for this post. Clearly I had done nothing illegal.

              WMF isn’t Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months’ rain fund in case something massive happens to current revenue.

              They can’t defend the contributors, but they can hire a union busting law firm to keep their staff in check. Got it.

              we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists

              We don’t know that because WMF hasn’t bothered to try defending contributors. Instead, they created a glide path for the pirates taking advantage of them.

              so they might as well make the exploiters pay for it to help maintain the commons’ infrastructure.

              I think it is interesting you say that WMF is being “exploited” yet you believe that they don’t have any standing for the damages they have experienced.

              • Aatube@lemmy.dbzer0.com
                link
                fedilink
                English
                arrow-up
                1
                ·
                1 day ago

                As I said, laws and ethics often get to different conclusions. They don’t have legal standing because there was no legal wrong, even if it is ethically abject.

                if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

                A cease and desist letter puts a person or business on notice that they’re engaging in an activity that violates your rights, and if they don’t stop, you’ll pursue legal action.

                In addition to saving time and money, a cease and desist letter is a good way to track communications you’ve sent to the other side. Sending a letter provides evidence that the party had notice of the wrongful behavior but continued to engage in it.

                https://www.nolo.com/legal-encyclopedia/what-is-a-cease-and-desist-letter.html

                Since Van Buren v. United States, scraping no longer falls under the powerful CFAA enforcement, and (I hope) the United States can no longer prosecute Aaron Swartz and seek seven years’ bars for scraping.

                but they can hire a union busting law firm

                Yeah that is absolutely horrible behavior; I agree with you entirely on the union parts which more people should know. (My post highlighting the hiring here has received surprisingly little attention, so I’m glad yours blew up.)

                However, that’s only at least $100k and at most a million, and to pursue a single entity. Mass-identifying ISPs from swarms of traffic-you-also-have-to-identify to which you write a letter each like Nintendo does with torrent swarms requires a lot more. It’s much easier and economical to simply put a rate limit on it and creating paid access on an isolated stack.

                hasn’t bothered to try

                Very false. Bot detection and blocking efforts have always been pretty documented (e.g. https://wikitech.wikimedia.org/wiki/Data_Platform/Data_Lake/Data_Issues/2026-06-10_Nov_2025_spike_in_bot_traffic , https://phabricator.wikimedia.org/project/board/5462/?filter=ve98KgxXcKmC , https://wikitech.wikimedia.org/wiki/Data_Platform/Data_Lake/Traffic/Bot_detection). They have only intensified since 2022 (https://diff.wikimedia.org/2026/03/26/quo-vadis-crawlers-progress-and-whats-next-on-safeguarding-our-infrastructure/image-1919/). Last year, they spent “~600–700 SRE FTE-hours” on reactive scraper defense by hand (“firefighting”).

                Instead, they created a glide path for the pirates taking advantage of them.

                Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

                • yoasif@fedia.ioOP
                  link
                  fedilink
                  arrow-up
                  1
                  ·
                  1 day ago

                  if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

                  This doesn’t have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.

                  I can’t imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.

                  We don’t know that because WMF hasn’t bothered to try defending contributors.

                  Very false. Bot detection and blocking efforts have always been pretty documented

                  I don’t see how bot detection shows that WMF is trying to defend contributors’ IP.

                  Instead, they created a glide path for the pirates taking advantage of them.

                  Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

                  Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies – if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it’s better to not be in the shadows, if you are trying to evade detection.

                  PS: Were they using the API or were they scraping? If they were using the API, why couldn’t WMF just revoke their keys?

    • TheTechnician27@lemmy.world
      link
      fedilink
      English
      arrow-up
      56
      arrow-down
      5
      ·
      edit-2
      4 days ago

      Nah, as someone who’s worked on Wikipedia for 10 years, this article and thread is a wild overreaction.

      The labor stuff I agree with, but in the grand scheme of things, it’s not that big a deal compared to the amount of work being done “below” the WMF within the projects themselves. It sucks, but this is a bump in the road. (Edit: I don’t mean to gloss over this, but I just assumed anyone here already knew the details; I should’ve linked to The Signpost article anyway.)

      On the other hand, the article treating Wikimedia Enterprise as a “betrayal” of its editors is patently absurd. It’s there so that companies and large institutions can access the data in a way that’s 1) minimally invasive to the project, 2) genenerating revenue for the project, and 3) timely for users; by contrast, without this sort of program, Wikimedia projects get hammered harder through unofficial means, make nothing from it, and plausibly get credited for outdated information. Material published to Wikimedia is CC BY-SA 4.0; thus, those current Enterprise customers have every right to use the material basically however they see fit regardless of Enterprise. And therefore make no mistake: they will do so in any legal way they can, which could include hammering Wikimedia or finding a way to bypass them nearly altogether.

      Wikimedia isn’t selling access to the material, because it’s literally nobody’s to sell; they’re selling access to stream the data on their servers which they host. And it’s not gatekeeping it from normal users either; it’s offered for free for what the Meta-Wiki accurarely calls “the vast majority of use cases” (they also give the Internet Archive free access).


      Source: ~50,000+ contributions to Wikimedia projects.


      Edit: Something I find peculiar in this article:

      One one side [sic], we have the Wikimedia Foundation, led by someone who hasn’t edited Wikipedia

      It strikes me that this entire article, this person never once (it’s bloviating, so maybe I missed it) mentions any work they’ve personally contributed to a Wikimedia project or even wikis generally. I can say outright as an actual volunteer who adores the other people who volunteer for the project that I could not care less if the CEO of the Wikimedia Foundation hasn’t contributed to a Wikimedia project a single time. That’d be cool, but heading a nonprofit is so divorced from the daily volunteer work that I don’t care, and I don’t think it has any bearing on her attempt at union-busting. (They also just say “edited Wikipedia”, which shows a level of project erasure actual editors wouldn’t make; Wikipedia has several sister projects, they’d be addressed collectively as “Wikimedia projects”, and veteran editors value their contributors just as highly if not moreso.)

      I’m not saying they’re not allowed to criticize anything they aren’t a part of, but I am saying “put up or shut up” when the author claims a project they could’ve contributed to for the past 25 years but ostensibly haven’t is under attack by ignorant/callous outsiders. More than almost anything, Wikimedia is excruciatingly open to work on if you really care (and this article shows they have the basic skills), so I expect an article this lengthy to disclose that when they start throwing other people under the bus for not having edited.


      Edit 2: So anyway, this was my essay about why you should check out Wiktionary and use it as your daily driver.

      • yoasif@fedia.ioOP
        link
        fedilink
        arrow-up
        8
        arrow-down
        7
        ·
        4 days ago

        Clearly, everyone realizes that the big tech AI pirates will scrape and stream the data - what I am objecting to is the response. When Google began to pirate Disney’s IP, Disney didn’t immediately offer them a data sharing deal (that also somehow doesn’t provide access to the IP) - they sued.

        The WMF has $296M in assets and they are grubbing for the pocket change that big tech throws at them for the fruits of the unpaid labor in Wikipedia. Why aren’t they suing for us? They are the stewards of the corpus.

        Has WMF even threatened the big tech owners with a good time, or were they simply salivating for the addition to their bottom line? Did they tell them to download the database? Did they attempt to ban their servers? Or did they simply provide big tech with privileged access to data they do not own?

        While clearly commentary from the community will be more interesting than from it outside of it, I don’t begrudge analysis or reporting from traditional media outlets - or do we want this all to be a private matter that 404 Media and the like don’t cover, allowing the theft (and union busting) to continue apace?

        In any case, I’ll show you mine.

        [Removed a reference to a different comment that I should have investigated more deeply.]

        • TheTechnician27@lemmy.world
          link
          fedilink
          English
          arrow-up
          10
          arrow-down
          5
          ·
          edit-2
          4 days ago

          Edit:

          [Edit:] PS: Now I understand why you attacked my post

          You know, I was polite about this in this comment, but now: screw you. My arguments above are why I attacked this garbage, bloviating post. I have more skin in this game in my pinky finger than you have in your entire body, so I have every right to an informed, detailed opinion that you don’t get to shut down by mining for a comment where I said “GPTs are useful to experts but problematic in the hands of the general public; also, I don’t personally use generative AI” – such a controversial statement. Really shows I’m sucking on the teat of Big AI that I don’t even use their product.


          Original:

          In any case, I’ll show you mine

          I appreciate that and your contributions (I would’ve believed you if you’d just said you had), and I just wish you’d established that before attacking someone else for not having volunteered contributions and tried to claim volunteers’ work is being “stolen”.

          Now with that: where on Earth was I “appealing to authority”? You’re just making up logical fallacies I didn’t even use. My links to The Signpost and the Meta-Wiki both have their own sources and valid explanations. As for “ad hom”, if you can attack people for being outsiders, I could absolutely question you for the same thing when you failed to assert you’d done anything within the project.

          • yoasif@fedia.ioOP
            link
            fedilink
            arrow-up
            6
            arrow-down
            6
            ·
            4 days ago

            As for “ad hom”, if you can attack people for being outsiders, I can absolutely question you for the same thing.

            She runs the foundation and is firing union workers. I’m just writing commentary. We are not the same.

            Volunteers’ work is stolen no matter if it is my own or others - me being an editor has no bearing on that.

            Now with that: where on Earth was I “appealing to authority”?

            Hmm: “Source: ~50,000+ contributions to Wikimedia projects.”

            • TheTechnician27@lemmy.world
              link
              fedilink
              English
              arrow-up
              8
              arrow-down
              2
              ·
              4 days ago

              Hmm: “Source: ~50,000+ contributions to Wikimedia projects.”

              My actual sources for statements of fact were already listed. It’s a shorthand for “This is my personal position as an experienced editor that this is an overreaction” that you’re warping in bad-faith.

              • ivanvector@piefed.ca
                link
                fedilink
                English
                arrow-up
                4
                arrow-down
                1
                ·
                4 days ago

                The volunteers’ work is in no way being stolen. All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can’t restrict access to any use complying with that license. Meaning that if AI companies want to scrape Wikipedia’s content for a compliant use, they will and they are fully legally permitted to.

                The Disney example is not apt: Disney produces content to make money, and reserves all rights to that content. It is not intended to be free nor to make money for third parties (absent a licensing agreement), and they were right to sue when AI companies misused their content. Wikipedia’s goals of creating a large collection of quality information and making it available to everyone for free are far from the same. The WMF does not own the content anyway, only the servers where the content resides.

                The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot. AI agents’ ridiculous server demands threatened to impede access for everyone else, and the deals they have made sidestep that impending problem while also generating some revenue for the WMF.

                The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content’s license.

                • yoasif@fedia.ioOP
                  link
                  fedilink
                  arrow-up
                  2
                  arrow-down
                  1
                  ·
                  4 days ago

                  All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can’t restrict access to any use complying with that license.

                  The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content’s license.

                  Why do you think the license allows for big tech to produce derivative works that are not licensed CC-BY-SA?

                  The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot.

                  If WMF has no right to decide who can access it and who cannot, how can they sell privileged access to who can access it? Frankly, that assertion fails on its face.

              • yoasif@fedia.ioOP
                link
                fedilink
                arrow-up
                2
                arrow-down
                1
                ·
                3 days ago

                My actual sources for statements of fact were already listed. It’s a shorthand for “This is my personal position as an experienced editor that this is an overreaction” that you’re warping in bad-faith.

                I have no idea why you think I am “warping” my reaction in bad faith - my reaction is based on the license text and what is written in the post. I also think it is ironic that you ask me to extend you grace in accepting your shorthand, and you clearly don’t bother to accept mine (that your statement was a reference to your own authority as an experienced editor).

                Material published to Wikimedia is CC BY-SA 4.0; thus, those current Enterprise customers have every right to use the material basically however they see fit regardless of Enterprise.

                You don’t actually tell us why this is the case - I argue that the companies are violating the license by not licensing their derivative works reciprocally - you don’t even bother to respond to that and just posit that they have “every right to use the material however they see fit”. Do you really believe that? Are the CC-BY-SA and GFDL licenses just completely worthless?

                Wikimedia isn’t selling access to the material, because it’s literally nobody’s to sell; they’re selling access to stream the data on their servers which they host.

                I’m not sure how much that matters. Would Warner Brothers not have an issue with me “streaming” access to their movies via my home server for payment?

                I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

                Enabling non-violating use cases don’t erase the violating ones.

    • webghost0101@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      31
      arrow-down
      3
      ·
      4 days ago

      “Maybe its time for wikipedia to die and for the copyleft corpus of data to evolve into something new” Was not on my bingocard and yet feels appropriate for the times.

      Information would not be lost but the public image of the web if Wikipedia stops existing would scatter overnight.

      • WhatAmLemmy@lemmy.world
        link
        fedilink
        English
        arrow-up
        39
        ·
        4 days ago

        This is all part of the plan for techno fascists. They need the Internet to fracture and become a place where everyone is trapped in a bubble of personalised propaganda that they control.

        That’s why they’re acquiring/consolidating and destroying all non-fascist sources across the media/web, after multiple decades of constructing pro-fascist alternatives.

    • Cris_Citrus (he/him <3)@piefed.zip
      link
      fedilink
      English
      arrow-up
      5
      ·
      edit-2
      3 days ago

      Support people you know. Support people in your community. Even when they have their issues you can help to them to help get them back on the right track- we don’t have the power to positively impact those who have harmful ideology when we aren’t in community with them

      The human beings around you could always use your support. If you don’t know where to start, start introducing yourself to your neighbors. Either when you see them out and about, or go knock on some doors.

  • Malyca@lemmy.zip
    link
    fedilink
    English
    arrow-up
    9
    ·
    3 days ago

    I’m not donating this year or ever again. I downloaded the whole thing in 2020 when the writing was on the wall, that’ll have to do.