IP blocks don't work, because they're using proxy networks so that you see an ip address 1 or 2 times within 10 minutes. They have effectively infinite ip addresses. (actually, looking at my data from today, I think this relationship holds over ~3 hours, where we're seeing ip address cardinality at about 1/2 of the hits.)
* Sometimes there's a pattern to the country. Oftentimes, not.
* User-agent, rotated between common, valid, current web browsers.
* Other headers, sec-*, accept, etc, generally valid and rotating.
* Bots will load the site to saturation in a denial of wallet attack.
The only thing that's specific is:
* urls have a pattern.
* it's obviously invalid traffic.
(non-bot traffic on my sites does not go from 0 to 200r/sec on the search interface in seconds. It does not go away that fast either)
Seeing the exact same thing on (somewhat high profile) open data sites I run.
The crawlers get stuck in a loop requesting the dataset listing page with every. single. combination. of. facets. At essentially as fast as it can be pumped out or blocked.
Had one bot super interested in a single organization to the tune of 1000 r/min for 24+ hours. Realistic rotating user agents, realistic sec-* headers, no ip address seen more than a couple times in 10 minutes. The only commonality was the route.
Yeah my search engine saw traffic of up to 160 queries per second the other day from some bot that was ostensibly searching for information on Jack Parsons. Just variations on the same query in different permutations of filters and site:-terms.
No I think this is the same case. It seems to just be following local links on the SERP. From a search for jack parsons, you can find hyperlinks to the sorts of requests it's making.
Once detected, don't block them because they'll just change strategies automatically, but you can toy with them, like returning a page full of random numbers instead of real data.
I’d love to, but I’m a bit limited in what I can do from a reputational damage POV. They’re my sites, in that I’m responsible, but they aren’t something where I can return incorrect responses.
I don’t think I’ve noticed gym socks smells, but I can certainly smell kerosene basically every time when outdoors at the airport, and often in the plane.
This is because it's an airport and there are tens of thousands of gallons of kerosene around, including several gallons per minute per running jet engine actively burning and turning into combustion products.
Even in a perfectly working bleed air system, the air comes from outside the plane, which on the tarmac of an airport, is god awful, shitty, and not at all healthy.
Lot's of other commenters talking about how the smell goes away at altitude.
Because it doesn't come from the packs (except in fume events), and what people are usually smelling is that airports have very bad air, and that air is much cleaner at 30k feet.
The miniscule oil vapor/droplets that are making their way into the bleed air system from engine compressor seals will still happen at 30k feet. If you don't smell something at all times on a plane, it cannot be coming from that.
This level of contamination would be like working at a gas station, which can be a health hazard. OSHA regulates air quality for employees working around hydrocarbons.
When I did it, I closed the door at nautical twilight. The key is that you need to close the door after the chickens roost but before the nocturnal predators come out.
Yes. We used to have these. And i turned it to most conservative twighlight.
It worked most of the times, until in winter, twice, some chickens were locked out. First time I saw it in time. Second time I only found feathers the next morning and one chicken less in the coop.
While on an especially cloudy (snowy) evenings, the door would be open when it was almost pitch dark out there and I caught a fox near already.
"The Real World" is just messier than a mathematical representation can be.
I use https://sunclock.net/ and have built a watchface like that for my fitbit like that. The mathematical model is perfect as guidance. But it's not reliable enough to bet the lives of my chickens on it.
Had something similar to this with the kids with Lego Duplo tracks, similar topology but the switches gave the system state. Going "backwards" through a switch set the train to return that direction when coming the other way.
So we had a goal to make the train do interesting behaviors, like cover the entire track autonomously, pushing the switches itself. The biggest run I remember was a 4 bit counter.
Kid is now entering his last year of CS + Maths degree, so I guess it checks out.
What’s more, they will scale up with increased resources on the site.
If you redline at 20 searches a sec, and put in 4 more workers, suddenly you’re serving 100r/sec to the bots, paying 5x for it, and your users are still seeing shit qos. I've seen multiple cores of nginx saturated just dealing with one dos/crawl run on a somewhat high profile site.
i would like to see a plot of the population of paleontologists starting from about 6-10 years after Jurassic Park was released. I suspect we'd see a bump starting around the time all the kids who saw Jurassic Park when it came out started graduating college.
We had Denver and Barney and The Land Before Time but kids suddenly memorizing all the latin names of each species was not a thing before Jurassic Park. (I think.)
reply