Fifty days of attacks on geheim.land, read in OpenCTI
I loaded fifty days of hostile traffic against this server into OpenCTI. This post uses what shows up there to explain what threat intelligence is, who knocks, how they do it, and what I take from it.
Between 17 August and 5 October 2026, the server behind geheim.land received hostile traffic from 33,473 IP addresses. They tried 226,378 passwords over SSH, sent 43,677 malicious web requests and opened 216,746 connections that the firewall refused. None of this is unusual. Every machine connected to the internet receives the same kind of traffic.
These numbers alone do not help much. What helps is knowing who is knocking, how they do it, and what to change on the server. Getting from the numbers to those answers is the job of cyber threat intelligence, or CTI. To show how it works, I took my own logs and loaded them into OpenCTI, an open source platform for managing threat intelligence.
From a log line to intelligence
A log line is raw data. One of mine says that 124.158.13.4 asked for /index.php?-d+allow_url_include…. Once it is put into shape it becomes information, with an address, a date and an attack technique. It becomes intelligence when I can act on it. This bot is looking for a known PHP flaw and the site has no PHP, so there is nothing to fear from it, but its signature is distinctive enough to block it everywhere.
The work has four steps. First I collect, here from the systemd journal for SSH and the firewall, from nginx for the web and from fail2ban for bans. Then I structure, by describing everything in a common format called STIX 2.1. Then I add context, by tying each address to a country, a hosting provider, a technique and a group. Last comes sharing, and deciding what to do.
The STIX file
I read 243 log files, 2,236,333 lines in all, from a copy of the logs and never from the originals. The output is a STIX 2.1 file, a JSON bundle of typed objects linked to each other. The version I load here is the one I call "essentiel". It keeps 5,855 addresses, those seen at least ten times and every one that attacked the web.
| STIX object | Count | What it holds |
|---|---|---|
ipv4-addr, ipv6-addr |
5,787 and 68 | the addresses that knocked |
autonomous-system |
1,070 | the provider or carrier of each address |
location |
116 | the country |
attack-pattern |
5 | the MITRE ATT&CK techniques seen |
intrusion-set |
157 | groups of addresses that look alike |
identity |
17 | geheim.land as the author, and 16 research scanners |
report |
1 | the report that holds everything |
relationship |
16,056 | the links (located at, related to, uses) |
Here is one address, with its description shortened.
{
"type": "ipv4-addr",
"value": "34.79.124.68",
"x_opencti_score": 60,
"x_opencti_labels": ["banni-fail2ban", "web"],
"x_opencti_description": "Vue du 15/09/2026 au 15/09/2026.\nWeb : 26 requêtes (fichiers sensibles (25 fois), logiciels absents) ; /.env., /.svn/wc.db, /.env;, …\nBanni par fail2ban : geheim-404-rafale.",
"belongs_to_refs": ["autonomous-system--3a77d5d6-7c1b-5fd0-8d43-959504e7196a"]
}
Because STIX is a standard, any tool that speaks it can read the file. OpenCTI imports it from Data › Import. Everything arrives inside one Report object, with a TLP:CLEAR marking that lets anyone read and share it, a note on the method that says what counts as hostile, and a list of limits. The file is written in French, which is why some of the screenshots below quote French text.
The dashboard
I built a dashboard in OpenCTI and restricted every widget to this data.

The labels are firewall (pare-feu), SSH, research scanner (scanner-recherche), web, and banned by fail2ban (banni-fail2ban). The donut counts labels rather than addresses, since one address can hit the firewall and also try SSH. The firewall comes first with 3,169 addresses, then SSH with 2,436, research scanners with 1,165, the web with 1,047 and fail2ban bans with 622.

The country list is where it is easiest to draw the wrong conclusion. The United States come far ahead with 2,180 addresses, followed by the United Kingdom (379), Germany (351) and the Netherlands (334). China is only fifth with 260 and France sixth with 233. Geolocation says where a machine sits. It says nothing about who drives it, and most of these machines are rented in the cloud. Google hosts 947 of these addresses, Censys 292, Akamai 274, UCloud in Hong Kong 221, Microsoft 188 and DigitalOcean 179. Renting a machine for an hour costs a few cents, so anyone in the world can knock from a data centre in Virginia.
Research scanners and the score
A good share of this traffic is hostile only in a loose sense. Palo Alto Networks with Cortex Xpanse, Censys, Driftnet, ONYPHE, Shadowserver and a few universities scan the whole internet to map it, and they say so through their reverse DNS, their HTTP signature or their published ranges. The file sets them apart as organisations, with the label scanner-recherche.

The score is there to tell them apart from the rest. Each address carries a score from 0 to 100. In this file, declared scanners get 10 (1,165 addresses), addresses seen mostly at the firewall get 30 (2,016) and SSH password guessers get 70 (1,959). Web addresses run from 50 to 80, and 80 is kept for the most aggressive ones, which try an exploit or dig through sensitive files in bulk (152).
A firewall or a SIEM that reads this data can block everything above 70 and ignore everything below 30. Without a score, the few real threats would be lost among thousands of harmless scans.
ATT&CK
Each address is linked to one or more techniques from MITRE ATT&CK, a catalogue that gives a stable name and identifier to each thing attackers do.

| Technique | What it is | Addresses |
|---|---|---|
| T1595.001 Scanning IP Blocks | sweeping addresses and ports | 4,490 |
| T1110.001 Password Guessing | guessing passwords | 2,036 |
| T1595.002 Vulnerability Scanning | looking for vulnerable software | 466 |
| T1595.003 Wordlist Scanning | looking for sensitive files | 382 |
| T1190 Exploit Public-Facing Application | trying an exploit | 172 |
My OpenCTI instance already had ATT&CK, imported by the MITRE connector. When my file brought in T1110.001, OpenCTI recognised it as the same object and merged the two. My technique then inherited everything MITRE knows about it, starting with the four mitigations it lists, among them multi-factor authentication and password policies. Of everything OpenCTI did with my file, this merge is what I found most useful.

The merge has a cost. MITRE's description was replaced by mine, "Devinette de mot de passe.", which is French for password guessing. When two sources describe the same object some fields overwrite others, so it is worth checking which ones win before importing.
The Knowledge tab lists who else uses the technique, 64 intrusion sets in all. 61 of them are my groups. The other three are APT28, APT29 and VOID MANTICORE, state-backed groups tracked by MITRE.

Sharing T1110.001 with APT28 says nothing about who my visitors are. Guessing an SSH password is within reach of any script. What I get from the shared technique is the vocabulary, and the list of defences MITRE has already written for it.
Groups
The file gathers addresses that look alike into 157 intrusion sets, using two criteria. 58 groups share the same tool, meaning the same paths requested, the same usernames tried or the same ports probed, from three addresses up. 99 groups share the same block, meaning the same /24 at the same provider, from five addresses up and twenty for the firewall. That gives names like GL outil ssh 49ffe8 for a tool and GL bloc pare-feu 162.216.150.0/24 for a firewall block.
Each group carries a confidence of 30 out of 100, which OpenCTI shows as doubtful. Every description opens with the same sentence, "Hypothèse de regroupement, pas une attribution", which means the group is a hypothesis and in no way an attribution.

The biggest one, GL outil ssh 49ffe8, is also the one that tells the most.

147 addresses tried exactly the same three usernames, root, ubuntu and geheim. The third one is the name of the domain, so the bot read the site's address and built an account name from it. Only 109 of the 147 were seen often enough to enter the "essentiel" file, which is why Figure 7 shows 109. Over the whole period, geheim was tried at least 5,009 times by 411 addresses. That makes it the fourth most tried username, after root (101,096), ubuntu (8,655) and admin (8,584). There is no account called geheim on this server.
The usernames also show what is being hunted. solana, solv, sol and firedancer go after validators of the Solana blockchain, and 195.178.110.30 alone tried 5,059 passwords, mostly on those accounts. wallet, crypto and bitcoin follow the same logic. claude comes up 546 times, I suppose because some people now create a user of that name for the AI assistant.
These totals come from the usernames listed for each address, and those lists are truncated, so the real numbers are higher.
Pivoting on a signature
The analysis really starts when I follow one clue from one address to the others. Take 124.158.13.4, with a score of 80.

It sent 46 requests between 26 August and 6 September, 44 of them exploitation attempts, and the paths are well known. /?%ADd+allow_url_include%3d1… goes for CVE-2024-4577, an argument injection in PHP-CGI on Windows, where the soft hyphen %AD is turned into a dash and lets the attacker pass options to PHP. /vendor/phpunit/phpunit/src/Util/PHP/eval-stdin.php goes for CVE-2017-9841, code execution through a copy of PHPUnit left behind in production.
Its HTTP signature is libredtail-http. The name points to RedTail, a cryptomining malware publicly documented since 2024 and known for exploiting this kind of flaw.
Searching the file for that signature, I find 44 addresses. 23 of them sit in two separate groups, GL outil web d4b481 (9 addresses, 46 paths) and GL outil web 7768d1 (14 addresses, 48 paths). The tool made two hypotheses of them only because their path lists differ by two lines. I put both groups side by side in an OpenCTI investigation graph.

A third orange group sits in the middle of the graph, GL outil ssh a8bf2a. Two of the 23 addresses, 95.85.245.227 and 138.197.189.197, also tried SSH passwords with the same three usernames (admin, root, user) as 28 other machines. So at least two of these bots look for PHP flaws on the web and guess passwords on SSH as well.
The techniques are the same, the paths are the same but for two, the signature is the same, and the addresses are spread over 11 countries and 19 providers, from DigitalOcean to Tencent by way of Oracle. To me it looks like one botnet seen from two angles. Deciding whether to merge the two hypotheses is my job, and if I do, the reason goes into the description of the group.
A fake Googlebot
The last case is 94.154.46.248.

It sent 4,425 requests between 6 and 30 September, 3,731 of them for sensitive files such as /.env, /wp-config.old, /.env.swp, /.aws/credentials and /settings.php, along with over a thousand other paths. Its HTTP signature claims to be Googlebot/2.1, Google's crawler.
The address belongs to Omegatech LTD, and Google has nothing to do with it. Nine addresses in the file call themselves Googlebot and none of them is at Google. Seven are at Omegatech, all in 94.154.46.0/24, one is at DEDIK SERVICES and one at Heymman Servers. The real Googlebot can be checked, since its reverse DNS ends in googlebot.com or google.com and Google publishes its address ranges. The disguise is meant to get past filters that let search engines in.
An HTTP signature is whatever the client chooses to send, so on its own it proves nothing. It has to be checked against what the attacker does not control, the provider, the reverse DNS and the published ranges.
What I take from it
On SSH, the mitigation MITRE attaches to T1110.001 fits a server like this one. Logging in should take a key and never a password, root should not log in directly, and no account should be named after the domain. With that in place, the 226,378 attempts have nothing to find.
On the web, nothing should sit at the root of the site, no .env, no .git, no backup. /.env was requested by 250 addresses in the file. The fail2ban rule geheim-404-rafale, which bans bursts of 404 errors, banned 103 of the addresses in the file.
I would rather block behaviours than addresses. A cloud address changes hands within hours, while a signature like libredtail-http, a burst of requests for /.env or a fake Googlebot can stay blocked for good. When an address does get blocked, it should have a high score and the block should expire after a few weeks.
Sharing comes last. The file is in STIX and marked TLP:CLEAR, so it can be passed on freely. Someone running OpenCTI could load it and compare it with their own logs, and an address seen by two servers is easier to judge than one seen by a single server. I have not put the file online for now.
Limits
Countries and providers come from DB-IP Lite, accurate to the month. An address can be shared by a cloud, a VPN or a mobile carrier, so behind a hostile IP there may be an innocent user. The logs show attempts and say nothing of their outcome. The groups are hypotheses at 30% confidence, and none of the group names here is an attribution. This file is the essential one, with 5,855 addresses out of 33,473.
Doing it yourself
OpenCTI runs with Docker Compose, from the OpenCTI-Platform/docker repository. Once the platform is up, you drop the STIX file into Data › Import and validate. The MITRE ATT&CK connector brings the catalogue. The dashboard and the investigation graph are then built in the interface, without writing any code.