IT administrators are struggling to deal with the ongoing fallout from the faulty CrowdStrike update. One spoke to The Register to share what it is like at the coalface.

Speaking on condition of anonymity, the administrator, who is responsible for a fleet of devices, many of which are used within warehouses, told us: “It is very disturbing that a single AV update can take down more machines than a global denial of service attack. I know some businesses that have hundreds of machines down. For me, it was about 25 percent of our PCs and 10 percent of servers.”

He isn’t alone. An administrator on Reddit said 40 percent of servers were affected, along with 70 percent of client computers stuck in a bootloop, or approximately 1,000 endpoints.

Sadly, for our administrator, things are less than ideal.

Another Redditor posted: "They sent us a patch but it required we boot into safe mode.

"We can’t boot into safe mode because our BitLocker keys are stored inside of a service that we can’t login to because our AD is down.

  • db0@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    158
    ·
    4 months ago

    Pity the administrators who dutifully kept a list of those keys on a secure server share, only to find that the server is also now showing a screen of baleful blue.

    Lol, can you imagine? It empathetically hurts me even thinking of this situation. Enter that brave hero who kept the fileshare decryption key in a local keepass :D

    • sugar_in_your_tea@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      92
      ·
      edit-2
      4 months ago

      That’s why the 3-2-1 rule exists:

      • 3 copies of everything on
      • 2 different forms of media with
      • 1 copy off site

      For something like keys, that means:

      1. secure server share
      2. server share backup at a different site
      3. physical copy (either USB, printed in a safe, etc)

      Any IT pro should be aware of this “rule.” Oh, and periodically test restoring from a backup to make sure the backup actually works.

      • IphtashuFitz@lemmy.world
        link
        fedilink
        English
        arrow-up
        19
        ·
        4 months ago

        We have a cron job that once a quarter files a ticket with whoever is on-call that week to test all our documented emergency access procedures to ensure they’re all working, accessible, up-to-date etc.

    • kescusay@lemmy.world
      link
      fedilink
      English
      arrow-up
      50
      ·
      4 months ago

      Seems like an argument for a heterogeneous environment, perhaps a solid and secure Linux server to host important keys like that.

        • Voroxpete@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          26
          ·
          4 months ago

          Their point is not that linux can’t fail, it’s that a mix of windows and linux is better than just one. That’s what “heterogeneous environment” means.

          You should think of your network environment like an ecosystem; monocultures are vulnerable to systemic failure. Diverse ecosystems are more resilient.

        • gnutrino@programming.dev
          link
          fedilink
          English
          arrow-up
          17
          ·
          4 months ago

          Sure but the chances of your Windows and Linux machines shitting the bed at the same time is less than if everything is running Windows. It’s exactly the same reason you keep a physical copy (which after all can break/burn down etc.) - more baskets to spread your eggs across.

          • pearsaltchocolatebar@discuss.online
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            4 months ago

            Very few businesses are going to spend the money running redundant infrastructure on two different operating systems. Most of them won’t even spend the money on a proper DR plan.

          • Avatar_of_Self@lemmy.world
            link
            fedilink
            English
            arrow-up
            1
            ·
            4 months ago

            Yes, but has it taken both OS’ out at the same time? It hasn’t but it could happen, however, the chances are even less. There’s obvious risk mitigation in mixing vendors in infrastructure for both hardware and software in the enterprise.

            If some critical services were lost in your enterprise last time until RH updated their kernel then you could have benefitted from running that service from Windows as well. Now the reverse is true. You could have another DC via Samba on Linux in your forest if you wanted to, in order to have an AD still for example. Same goes for file share servers, intermediary certificate servers (hopefully your Root CA is not always on the network) and pretty much most critical services.

            Most enterprises run a lot of services off of a hypervisor and have overhead to scale (or they are already in a sinking ship), so you can just spin up VMs to do that. It isn’t as if it is unreasonably labor intensive compared to other similar risk mitigation implementations. Any sane CCB (obviously there are edge cases but we are talking in general here) will even let you get away without a vendor support contract for those, since they are just for emergency redundancy and not anywhere near critical unless the critical services have already shit the bed.

  • catloaf@lemm.ee
    link
    fedilink
    English
    arrow-up
    128
    ·
    4 months ago

    We can’t boot into safe mode because our BitLocker keys are stored inside of a service that we can’t login to because our AD is down.

    Someone never tested their DR plans, if they even have them. Generally locking your keys inside the car is not a good idea.

    • jet@hackertalks.com
      link
      fedilink
      English
      arrow-up
      36
      ·
      4 months ago

      The good news is! This is a shake out test and they’re going to update those playbooks

      • Justin@lemmy.jlh.name
        link
        fedilink
        English
        arrow-up
        31
        arrow-down
        1
        ·
        4 months ago

        Sysadmins are lucky it wasn’t malware this time. Next time could be a lot worse than just a kernel driver with a crash bug.

        3rd party companies really shouldn’t have access to ship out kernel drivers to millions of computers like this.

      • Evotech@lemmy.world
        link
        fedilink
        English
        arrow-up
        11
        ·
        4 months ago

        The bad news is that the next incident will be something else they haven’t thought about

      • ɔiƚoxɘup@infosec.pub
        link
        fedilink
        English
        arrow-up
        9
        ·
        4 months ago

        I wish you were right. I really wish you were. I don’t think you are. I’m not trying to be a contrarian but I don’t think for a large number of organizations that this is the case.

        For what it’s worth I truly hope that I’m 100% incorrect and everybody learns from this bullshit but that may not be the case.

    • Zron@lemmy.world
      link
      fedilink
      English
      arrow-up
      18
      ·
      4 months ago

      I remember a few career changes ago, I was a back room kid working for an MSP.

      One day I get an email to build a computer for the company, cheap as hell. Basically just enough to boot Windows 7.

      I was to build it, put it online long enough to get all of the drivers installed, and then set it up in the server room, as physically far away from any network ports as possible. IIRC I was even given an IO shield that physically covered the network port for after it updated.

      It was our air-gapped encryption key backup.

      I feel like that shitty company was somehow prepared for this better than some of these companies today. In fact, I wonder if that computer is still running somewhere and just saved someone’s ass.

    • ripcord@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      4 months ago

      They also don’t seem to have a process for testing updates like these…?

      This seems like showing some really shitty testing practices at a ton of IT departments.

        • ripcord@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          4 months ago

          I’ve heard differently. But if it’s true, that should have been a non-starter for the product for exactly reasons like this. This is basic stuff.

          • Entropywins@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            4 months ago

            Companies use crowdstrike so they don’t need internal cybersecurity. Not having automatic updates for new cyber threats sorta defeats the purpose of outsourcing cybersecurity.

            • hangonasecond@lemmy.world
              link
              fedilink
              English
              arrow-up
              2
              ·
              4 months ago

              Automatic updates should still have risk mitigation in place, and the outage didn’t only affect small businesses with no cyber security capability. Outsourcing does not mean closing your eyes and letting the third party do whatever they want.

              • kent_eh@lemmy.ca
                link
                fedilink
                English
                arrow-up
                3
                arrow-down
                1
                ·
                4 months ago

                Outsourcing does not mean closing your eyes and letting the third party do whatever they want.

                It shouldn’t, but when the decisions are made by bean counters and not people with security knowledge things like this can easily (and frequently) happen.

            • ripcord@lemmy.world
              link
              fedilink
              English
              arrow-up
              2
              ·
              4 months ago

              Not bothering doing basic, minimal testing - and other mitigation processes - before rolling out updates is absolutely terrible policy.

      • catloaf@lemm.ee
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 months ago

        Unfortunately, the pace of attack development doesn’t really give much time for testing.

          • TonyOstrich@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            ·
            4 months ago

            I was just thinking about something similar. I can understand wanting to get a security update as quickly as possible, but it still seems like some kind of rolling update could have mitigated something like this. When I say rolling, I mean for example split all of your customers into 24 groups and push the update once an hour to another group. If it causes a massive fuck up it’s only some or most, but not all.

    • JasonDJ@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      4 months ago

      I get storing bitlocker keys in AD, but as a net admin and not a server admin…what do you do with the DCs keys? USB storage in a sealed envelope in a safe (or at worst, locked file cabinet drawer in the IT managers office)?

      Or do people forego running bitlocker on servers since encrypting data-at-rest can be compensated by physical security in the data center?

      Or DCs run on SEDs?

      • catloaf@lemm.ee
        link
        fedilink
        English
        arrow-up
        4
        ·
        4 months ago

        When I set it up at one company, the recovery keys were printed out and kept separately.

  • gravitas_deficiency@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    92
    ·
    4 months ago

    Lmao this is incredible

    Another Redditor posted: "They sent us a patch but it required we boot into safe mode.

    "We can’t boot into safe mode because our BitLocker keys are stored inside of a service that we can’t login to because our AD is down.

    “Most of our comms are down, most execs’ laptops are in infinite bsod boot loops, engineers can’t get access to credentials to servers.”

    N.B.: Reddit link is from the source

    I hope a lot of c-suites get fired for this. But I’m pretty sure they won’t be.

    • MagicShel@programming.dev
      link
      fedilink
      English
      arrow-up
      67
      ·
      4 months ago

      C-suites fired? That’s the funniest thing I’ve heard yet today. They aren’t getting fired - they are their own ass-coverage. How can they be to blame when all these other companies were hit as well?

      I guess this is a good week for me to still be laid off.

    • Codex@lemmy.world
      link
      fedilink
      English
      arrow-up
      53
      ·
      4 months ago

      Our administrator is understandably a little bitter about the whole experience as it has unfolded, saying, "We were forced to switch from the perfectly good ESET solution which we have used for years by our central IT team last year.

      Sounds like a lot of architects and admins are going to get thrown under the bus for this one.

      “Yes, we ordered you to cut costs in impossible ways, but we never told you specifically to centralize everything with a third party, that was just the only financially acceptable solution that we would approve. This is still your fault, so we’re firing the entire IT department and replacing them with an AI managed by a company in Sri Lanka.”

      • Evotech@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        4 months ago

        Stupid argument though, honestly just chance that crowdstrike was the vendor to shit the bed. Might aswell have been set. You should still have procedures for this

  • MrNesser@lemmy.world
    link
    fedilink
    English
    arrow-up
    93
    arrow-down
    4
    ·
    4 months ago

    Lemmy appears to be weathering the storm quite well…

    …probably runs on linux

    • cygnus@lemmy.ca
      link
      fedilink
      English
      arrow-up
      76
      arrow-down
      1
      ·
      edit-2
      4 months ago

      The overwhelming majority of webservers run Linux (it’s not even close, like high 90 percent range) Edit: Upon double-checking it’s more like mid-80s, but the point stands

    • RBG@discuss.tchncs.de
      link
      fedilink
      English
      arrow-up
      56
      ·
      4 months ago

      It runs on hundreds of servers. If any of them ran windows they might be out but unless you got an account on them you’d be fine with the rest. That’s the whole point of federation.

    • Bilb!@lem.monster
      link
      fedilink
      English
      arrow-up
      5
      ·
      4 months ago

      I wonder if any Lemmy servers run on Windows without WSL. I can’t think of any hard dependencies on Linux, so it should be possible.

  • Boozilla@lemmy.world
    link
    fedilink
    English
    arrow-up
    61
    ·
    edit-2
    4 months ago

    If you have EC2 instances running Windows on AWS, here is a trick that works in many (not all) cases. It has recovered a few instances for us:

    • Shut down the affected instance.
    • Detach the boot volume.
    • Move the boot volume (attach) to a working instance in the same region (us-east-1a or whatever).
    • Remove the file(s) recommended by Crowdstrike:
    • Navigate to the C:\Windows\System32\drivers\CrowdStrike directory
    • Locate the file(s) matching “C-00000291*.sys”, and delete them (unless they have already been fixed by Crowdstrike).
    • Detach and move the volume back over to original instance (attach)
    • Boot original instance

    Alternatively, you can restore from a snapshot prior to when the bad update went out from Crowdstrike. But that is not always ideal.

    • Defaced@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      4 months ago

      A word of caution, I’ve done this over a dozen times today and I did have one server where the bootloader was wiped after I attached it to another EC2. Always make a snapshot before doing the work just in case.

  • Max-P@lemmy.max-p.me
    link
    fedilink
    English
    arrow-up
    48
    arrow-down
    1
    ·
    4 months ago

    This is why every machine I manage has a second boot option to download a small recovery image off the Internet and phone home with a shell. And a copy of it on a cheap USB stick.

    Worst case I can boot the Windows install in a VM with the real disk, do the maintenance remotely. I can reinstall the whole thing remotely. Just need the user to mash F12 during boot and select the recovery environment, possibly input WiFi credentials if not wired.

    I feel like this should be standard if you have a lot of remote machines in the field.

    • corsicanguppy@lemmy.ca
      link
      fedilink
      English
      arrow-up
      20
      ·
      4 months ago

      This is why every machine I manage has a second boot option to download a small recovery image off the Internet and phone home with a shell. And a copy of it on a cheap USB stick.

      You’re fucking killing it. Stay awesome.

      Also gist this up pls. Thanks.

      • Max-P@lemmy.max-p.me
        link
        fedilink
        English
        arrow-up
        18
        ·
        4 months ago

        I wish it was more shareable, but it’s also not as magic as it sounds.

        Fundamentally it’s just a Linux install with some heavy customizations so that it does one thing only: boot Linux, and just enough prompts to get it online so that the VPN works, and download the root image into RAM that it boots into so I can SSH into the box, and then a bunch of Linux tools for me to use so I can reimage from there, or run a QEMU with the physical disk passed through so I can VNC into an install even if it BSOD.

        It’s a Linux UKI (combined kernel+initramfs into a simple EFI file the firmware can boot directly without a bootloader), but you can just as easily get away with a hidden Debian install or whatever. Can even be a second Windows install if that’s your thing. The reason I went this particular route is I don’t have to update it since it downloads it on the fly, much like the Mac recovery. And it runs entirely in RAM afrerwards so I can safely do whatever is needed with the disk.

        • flop_leash_973@lemmy.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          4 months ago

          I dream of working somewhere where this kind of effort is appreciated enough to motivate me to put in the effort of actually doing it.

          • Max-P@lemmy.max-p.me
            link
            fedilink
            English
            arrow-up
            1
            ·
            4 months ago

            I wish too, it’s only deployed for family and family businesses because I’m a couple thousand miles away from them. I cobbled this together for the explicit purpose of being able to reinstall Windows remotely. It works wonderfully though!

            My real job is DevOps and 100% Linux, and most of the cloud servers are disposable and can be simply be rebuilt at the push of a button in some dashboard.

    • person420@lemmynsfw.com
      link
      fedilink
      English
      arrow-up
      8
      ·
      4 months ago

      Just need the user to mash F12 during boot and select the recovery environment, possibly input WiFi credentials if not wired

      In theory that sounds great, now just do it 1000+ times while your phone is ringing off the hook and you’re working with some of the most tech illiterate people in your org.

    • douglasg14b@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      arrow-down
      1
      ·
      edit-2
      4 months ago

      Sounds like a nightmare for security, and a dream for attackers.

      More companies need to do this, solid job security.

      • Max-P@lemmy.max-p.me
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 months ago

        You can sign the whole thing, it’s not like you have to turn off secure boot and just drop the user to a root shell. There’s nothing to be gained from it, especially if you have physical access to the machine.

  • Buffalox@lemmy.world
    link
    fedilink
    English
    arrow-up
    54
    arrow-down
    10
    ·
    edit-2
    4 months ago

    At least no mission critical services were hit, because nobody would run mission critical services in Windows, right?

    RIGHT??

    • Hotzilla@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      17
      ·
      4 months ago

      Issue is not just on servers, but endpoints also. Servers are something that you can relatively easily fix, because they are either virtualized or physically in same location.

      But endpoints you might have thousand physical locations, and IT need to visit all of them (POS, info/commercial displays, IoT sensors etc.).

      • Miaou@jlai.lu
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        1
        ·
        4 months ago

        Parent comment applies even more so to such endpoints imo

    • CaptPretentious@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      4 months ago

      I’m the corporate world, very much Windows gets used. I know Lemmy likes a circle jerk around Linux. But in the corporate world you find various OS’s for both desktop and servers. I had to support several different OS’s and developed only for two. They all suck in different ways there are no clear winners.

      • Dark Arc@social.packetloss.gg
        link
        fedilink
        English
        arrow-up
        0
        arrow-down
        1
        ·
        edit-2
        4 months ago

        It’s not just a circle jerk in this case. Windows is dominant for desktop usage but Linux has like 90% of the server market and is used for basically all new server projects.

        Paying for Windows licensing when it doesn’t benefit you, it’s silly, and that’s been realized for years.

    • kent_eh@lemmy.ca
      link
      fedilink
      English
      arrow-up
      5
      ·
      4 months ago

      My former employer had a bunch of windows servers providing remote desktops for us to access some proprietary (and often legacy) mission critical software.

      Part of the security policy was that any machines in the possession of end users were assumed to be untrustworthy, so they kept the applications locked down on the servers.

    • stoly@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      6
      ·
      4 months ago

      I can’t imagine how much work it would be to migrate all your services onto Linux. The problem was people adopting windows in the first place.

      • douglasg14b@lemmy.world
        link
        fedilink
        English
        arrow-up
        9
        arrow-down
        4
        ·
        4 months ago

        I love the Linux bros coming out of the woodwork on this one when this could have very well have been Linux on the receiving end of this shit show. Given that it’s a kernal level software issue, and not necessarily an OS one.

        It’s largely infeasible to use Linux for many, most, of these endpoints. But facts are hard.

        • save_the_humans@leminal.space
          link
          fedilink
          English
          arrow-up
          5
          ·
          edit-2
          4 months ago

          Hey man, let us have this one. Any immutable/atomic distribution could have either prevented this or easily rolled back the update. Not to mention a Linux offering by something like Red Hat, for example, wouldnt recommend installing closed source third party kernel modules for exactly this reason. Not sure about the feasibility of these endpoints, but the way things are generally done on, and the philosophy of, Linux could very well have avoided this catastrophe.

        • jabjoe@feddit.uk
          link
          fedilink
          English
          arrow-up
          3
          ·
          4 months ago

          The is no single Linux. It’s not a monoculture like that. There are many distros with different build options, different configurations and different components.

          Also culture is different. Very few Linux admins would be happy putting in a closed blob kernel driver for anything. In Windows world that’s the norm, but not Linux.

          What’s just happened to Windows world would be harder in Linux world. At worse, one distros rolls out a killer update. Some distros would just reboot to the previous kernel.

        • flop_leash_973@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          4
          ·
          edit-2
          4 months ago

          They are just butt hurt that this whole thing really shines a light on how inaccurate the line of “the world runs on Linux” truly is.

          The world runs on a lot of different things for different reasons and that does not fit nicely into their Richard Stallman like world view.

          • pathief@lemmy.world
            link
            fedilink
            English
            arrow-up
            1
            ·
            4 months ago

            Just to clarify: the world runs in linux servers. The market share for the non-server market is abysmal.

            • jabjoe@feddit.uk
              link
              fedilink
              English
              arrow-up
              3
              ·
              4 months ago

              Except lots of IoT things, router, etc. Also Cromebooks and Steamdecks. And us GNU/Linux people. Android is Linux, just not GNU/Linux. Really isn’t just servers.

              • pathief@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                ·
                edit-2
                4 months ago

                I don’t know too much about IoT but I wouldn’t say linux runs the world in any of the other markets you mentioned.

                I would say while technically Android uses a modified linux kernel, you can’t put it under the same umbrella.

                Either way I don’t want to get too much into these technicalities. I was simply trying to say that Linux is king on servers, not really on the market where all this crazyness happened.

    • ɔiƚoxɘup@infosec.pub
      link
      fedilink
      English
      arrow-up
      4
      ·
      4 months ago

      I’m in. This world desperately needs an information workers union. Someone to cover those poor fuckers in the help desk and desktop support as well as the engineers and architects that keep all of this shit running.

      Those of us that aren’t underpaid are treated poorly. Today is what it looks like if everybody strikes at once.

      • slacktoid@lemmy.ml
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 months ago

        This dude here coming in hot with a name, Information Workers Union (IWU). Love it

        Soo are you gonna create the community or am I?

  • Hurculina Drubman@lemm.ee
    link
    fedilink
    English
    arrow-up
    5
    arrow-down
    1
    ·
    4 months ago

    I got super lucky. got paid for my car just before the dealership systems went down, got my return flight 2 days before this shit started.

    • catloaf@lemm.ee
      link
      fedilink
      English
      arrow-up
      18
      ·
      4 months ago

      Because that’s where filesystem access lives? AV wouldn’t do very much good if it could only run from userspace.

      • Blaster M@lemmy.world
        link
        fedilink
        English
        arrow-up
        16
        arrow-down
        4
        ·
        4 months ago

        Pretending linux privelege escalation doesn’t exist… to fight something that gets root you have to be able to fight at the root level, or the root access malware can simply nuke the av from userland.

        • Justin@lemmy.jlh.name
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          2
          ·
          edit-2
          4 months ago

          Or you could just use kernel namespaces, SELinux, Systemd sandboxing, etc. There is zero need to run in ring 0 for security reasons.

          Also, privilege escalation is a lot rarer on Linux than it is on Windows.

  • scottywh@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    4
    ·
    4 months ago

    If it only impacts a percentage of your machines then there was a problem in the deployment strategy or the solution wasn’t worthwhile to begin with.

    • Phoenixz@lemmy.ca
      link
      fedilink
      English
      arrow-up
      0
      arrow-down
      7
      ·
      4 months ago

      … So your point was that it would have been better if everything went down?

      There are plentiful reasons why deployments are done in parts, and I’m guessing that after today strategies will change to apply updates in groups to avoid everything going down.

      Also, dear God, stop using windows as a server, or even a client for that matter. If you’re paying actual money to get this shit then the results are on you.