TL;DR: In the last year, the Wikimedia Foundation has fired several union organizers, including those that worked on the Community Tech team - a team dedicated to building features for the volunteer community that edits Wikipedia.

As the Wiki Workers Union tries to get the Wikimedia Foundation to recognize their union, it is worth remembering that this is not the first time that the Foundation has worked against the community.

Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus - a body of work that the Foundation does not own.

Movement volunteer communities contributed to Wikipedia under copyleft licenses - licenses that work to ensure that the work remains free (as in speech). The big tech AI companies do not license derivative works under copyleft licenses and often do not even attribute where the works came from.

This means that volunteers are working for big tech for free, and the Wikimedia Foundation is selling privileged access to that free labor.

It is against that backdrop that the current unionization struggle unfolds.

  • yoasif@fedia.ioOP
    link
    fedilink
    arrow-up
    1
    ·
    5 days ago

    Send one to whom? A WHOIS of the Brazilian IP that turns out to be a residential proxy? Anonymous scraping for LLMs is a problem all over the Internet without a solution (save Anubis). There’s a reason all the big companies have went for Cloudflare instead of any lawsuits, which you’d hear in the news.

    I know individuals that have been unmasked in torrent swarms and have had their ISP cancel their service due to that. The idea that Wikimedia’s hands are powerless to send a Cease and Desist to ISPs to warn and ban their customers for scraping is hilarious.

    Again, they have $296M in reserves. WMF can send a letter.

    They’re not even sure if they have the standing; they’ve looked at that question

    But that isn’t what you linked to says; the word standing doesn’t appear in the text, nor do they seem to explore that. Thanks for the reference, but it doesn’t actually support your argument.

    I am simply representing the perspective of contributors as a signatory myself, and as a developer who makes API calls to Wikipedia. “Our content is always free to use, but our infrastructure is not” sums it up nicely and Wikimedia Enterprise keeps the free infrastructure good.

    As I responded to another commenter:

    I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

    Enabling non-violating use cases don’t erase the violating ones.

    • Aatube@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      1
      ·
      5 days ago

      individuals that have been unmasked in torrent swarms

      WMF isn’t Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months’ rain fund in case something massive happens to current revenue.

      But that isn’t what you linked to says; the word standing doesn’t appear in the text

      “Standing” means you have been harmed by an illegal act. You need standing to cease and desist or sue. To have legal standing for a letter, what you’re asking to cease and desist needs to be illegal. What I linked to is that WMF counsel believes LLMs likely do not violate the content licenses of their training data. Thus, there is no legal injury and no standing.

      Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally

      see what I quoted. If we go by ethics (which I do prefer) instead of the law, we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists—in fact, they have been happening despite Enterprise (although at a reduced rate) to the point where rate limits have been enacted for the first time in history—so they might as well make the exploiters pay for it to help maintain the commons’ infrastructure.

      • yoasif@fedia.ioOP
        link
        fedilink
        arrow-up
        1
        ·
        5 days ago

        To have legal standing for a letter, what you’re asking to cease and desist needs to be illegal.

        FWIW, that isn’t true - I didn’t post the letter, but I got a C&D from SoFi for this post. Clearly I had done nothing illegal.

        WMF isn’t Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months’ rain fund in case something massive happens to current revenue.

        They can’t defend the contributors, but they can hire a union busting law firm to keep their staff in check. Got it.

        we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists

        We don’t know that because WMF hasn’t bothered to try defending contributors. Instead, they created a glide path for the pirates taking advantage of them.

        so they might as well make the exploiters pay for it to help maintain the commons’ infrastructure.

        I think it is interesting you say that WMF is being “exploited” yet you believe that they don’t have any standing for the damages they have experienced.

        • Aatube@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          1
          ·
          5 days ago

          As I said, laws and ethics often get to different conclusions. They don’t have legal standing because there was no legal wrong, even if it is ethically abject.

          if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

          A cease and desist letter puts a person or business on notice that they’re engaging in an activity that violates your rights, and if they don’t stop, you’ll pursue legal action.

          In addition to saving time and money, a cease and desist letter is a good way to track communications you’ve sent to the other side. Sending a letter provides evidence that the party had notice of the wrongful behavior but continued to engage in it.

          https://www.nolo.com/legal-encyclopedia/what-is-a-cease-and-desist-letter.html

          Since Van Buren v. United States, scraping no longer falls under the powerful CFAA enforcement, and (I hope) the United States can no longer prosecute Aaron Swartz and seek seven years’ bars for scraping.

          but they can hire a union busting law firm

          Yeah that is absolutely horrible behavior; I agree with you entirely on the union parts which more people should know. (My post highlighting the hiring here has received surprisingly little attention, so I’m glad yours blew up.)

          However, that’s only at least $100k and at most a million, and to pursue a single entity. Mass-identifying ISPs from swarms of traffic-you-also-have-to-identify to which you write a letter each like Nintendo does with torrent swarms requires a lot more. It’s much easier and economical to simply put a rate limit on it and creating paid access on an isolated stack.

          hasn’t bothered to try

          Very false. Bot detection and blocking efforts have always been pretty documented (e.g. https://wikitech.wikimedia.org/wiki/Data_Platform/Data_Lake/Data_Issues/2026-06-10_Nov_2025_spike_in_bot_traffic , https://phabricator.wikimedia.org/project/board/5462/?filter=ve98KgxXcKmC , https://wikitech.wikimedia.org/wiki/Data_Platform/Data_Lake/Traffic/Bot_detection). They have only intensified since 2022 (https://diff.wikimedia.org/2026/03/26/quo-vadis-crawlers-progress-and-whats-next-on-safeguarding-our-infrastructure/image-1919/). Last year, they spent “~600–700 SRE FTE-hours” on reactive scraper defense by hand (“firefighting”).

          Instead, they created a glide path for the pirates taking advantage of them.

          Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

          • yoasif@fedia.ioOP
            link
            fedilink
            arrow-up
            1
            ·
            5 days ago

            if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

            This doesn’t have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.

            I can’t imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.

            We don’t know that because WMF hasn’t bothered to try defending contributors.

            Very false. Bot detection and blocking efforts have always been pretty documented

            I don’t see how bot detection shows that WMF is trying to defend contributors’ IP.

            Instead, they created a glide path for the pirates taking advantage of them.

            Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

            Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies – if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it’s better to not be in the shadows, if you are trying to evade detection.

            PS: Were they using the API or were they scraping? If they were using the API, why couldn’t WMF just revoke their keys?

            • Aatube@lemmy.dbzer0.com
              link
              fedilink
              English
              arrow-up
              1
              ·
              5 days ago

              Fair point on using ToS.

              Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies

              It is extremely cheap to rely on residential proxies and certainly magnitudes cheaper than Enterprise. The free infrastructure is absolutely not shoddy and I’m not sure why you think it is, though making all scrapers use it would probably make it shoddy.

              If they were using the API, why couldn’t WMF just revoke their keys?

              The API has no keys. It’s commons public access. If you want to revoke access, you’d have to identify the IPs and manually block them.

              You’re also still assuming this is violating intellectual property. Customers of Wikimedia Enterprise are forced to sign Terms of Use (accessible at https://dashboard.enterprise.wikimedia.com/signup/) that legally bind them to respect the license of the content even though, as you point out, they are already bound to respect the license. To maintain access to Enterprise, you must not violate that license, and with that the WMF can send DMCA C&Ds.

              Since Enterprise does have API keys, it in fact gives WMF control over exploitation of it enforceable in the way you mention. Enterprise is what defends our copyright.

              • yoasif@fedia.ioOP
                link
                fedilink
                arrow-up
                1
                ·
                4 days ago

                Enterprise is what defends our copyright.

                But it clearly isn’t. Hence my mention of it in my post and why I see it as a betrayal.

                • Aatube@lemmy.dbzer0.com
                  link
                  fedilink
                  English
                  arrow-up
                  0
                  ·
                  4 days ago

                  would you kindly respond to what I said :)? restating a conclusion i’ve responded to is unfortunately not much for me to go off of

                  • yoasif@fedia.ioOP
                    link
                    fedilink
                    arrow-up
                    1
                    ·
                    4 days ago

                    Yes, I am assuming that this violates the Wikipedia IP, I explained that in my post. I’m not going to re-litigate that since you haven’t shown how I am wrong about that.

                    Given that, a C&D for discontinuing scraping, followed by a suit to desist any further scraping would seem to be in order.