Moderator's introduction
You can put the tincture over there, please. And we move on to the final talk for today. After that, only the roast is left. But what's the topic of the talk? Let's muse a little, now that the day is ending. The internet, in itself, is a very interesting place, because sometimes you find things there that you weren't even looking for. Not that long ago, about three years back, a story went around the web that at the bottom of a kindergarten's website someone sold illegal substances. I don't know if you saw it, but there was such a story. So on the first page, kittens, and on the second, a whole drug-dealing network.
Our next speaker, Igor Bederov, will tell us how cases of this kind get unravelled. And he'll do it right now. Igor, the floor is yours. The microphone and the clicker, please.
Talk and Q&A
Thank you very much, everything's working. Right, I've gone back. Yes, once more? Yep, once more. —
— And once more, yes. Thank you all for coming. And indeed, well, from forensics let's try to dive into OSINT. Competitive intelligence: collecting, researching public-source information. It's also very important for us. We've already discussed the prospects of forensics with colleagues in the hallway, especially the prospects, given that, possibly, in the upcoming iPhones, and other mobile phones too, they may drop the USB Type-C port, the charging port, and how forensics is going to develop at all, when there's simply nothing left to plug forensic software and hardware into. So, data analysis is probably going to develop as well. And although at our previous meetings we talked about Telegram users, about Telegram channels, today there seems to be, at least, as I've been told, a request from the audience for research into websites.
The topic seems perfectly simple and obvious. Websites have certain users we're interested in, and our task is to identify the people involved in owning this site, administering it, or developing it. So, the three roles that appear in our investigations are the owner, the administrator and the developer. The one who owns the domain name or the hosting, the one who administers it, communicates with users and posts this or that content, and the one who develops the site's engine, hooks various technologies up to it and administers all of that. Where to start? Let's start with the simplest thing, a pile of useful sources. You can take a photo, or you can not take a photo. This slide will come up again at the end.
My colleagues and I went to the trouble, specifically for everyone who does research and investigations, of creating our own build of the Opera browser, a portable browser that runs from a USB stick and stores your authorized sessions on that same USB stick. So you can take your workstation with you, and in this browser, besides a number of privacy settings, there is also a pile of useful sources, including for researching websites, and for other kinds of research that may be useful to you. Download it via the link, and if someone doesn't like Opera, there's an option to load all the useful sources into any other browser as an HTML file.
So, let's go. A web resource is, first and foremost, a domain. What we see in our browser is its domain name, the one we navigate to. vk.com, gosuslugi.ru, or some other one. All these domain names must be registered. I won't go into the details about the international organizations, about ICANN and so on. A domain name gets registered, and the information about that domain name is stored in various WHOIS services. Everything seems great, wonderful, but naturally we run more and more into domain names being registered extremely sloppily, the owners' details are not verified, not checked, and that poses a certain problem for us.
WHOIS data becomes unreliable, and with the arrival of a thing like GDPR and the restriction of personal data in WHOIS, we've come up against the fact that it says a private person is the domain owner, and what to do with that, we don't fully know either. But when we do OSINT, we understand that all of this can exist in historical retrospect, so when we need to get WHOIS data for old sites, we can dig into the WHOIS data archives, which are kept by a large number of services listed here on the screen. Yes, GDPR has limited the amount of personal data returned in WHOIS, but in WHOIS archives it may have been preserved, and we have to check that data in order to find out the name or the company that owns the domain name.
The next thing is a bit off-topic, but since in August the Russian Supreme Court issued a ruling for business entities, and I think there are probably security service representatives, or future security service representatives, here, obliging commercial entities to detect typosquatting on their own, and online fraud, I can straight away suggest a few simple and obvious resources that let you monitor the appearance of domain names that resemble your organization's domain name. And also monitor the leaks that happen involving your domain. They do it for free. So on one side we have DNSTwister, dnstwist and the like, which let you find domain names with a similar spelling, see if they're active or not, if a site has appeared there, whether email has appeared behind that resource. And IntelX, Have I Been Pwned are services that let you find out if there have been leaks on your domain, and thus protect the organization.
The next thing, after we've talked about the domain name, is hosting, the physical location of our site. Our site is images, texts, some volume of information; it has to physically reside somewhere. It resides on an external server, which is called hosting. To figure out which hosting our site is using, there is also a large number of services, starting with a plain ping, which you can do from the operating system, and ending with a pile of external services.
Now, the most important and central thing, probably, about web hosting, is Cloudflare protection. Originally it was created so that we could minimize DDoS attacks, but in practice Cloudflare is actively used by offenders to hide the actual location of their site. And a frequently asked question is how the Cloudflare protection could possibly be stripped away to understand where our site is physically located. And you can't always strip it away, but to some extent you can, through external services that may have indexed our web resource before the Cloudflare protection was put in place. That's URLScan, that's VirusTotal, that's various leaks of Cloudflare itself that are out there on the internet, that's DNS data analysis, that's analysis of all the other systems and technologies present on our site that verify the site in Yandex, Google and other systems, load new technologies, payment acquiring and the like.
And finally, it's looking for reuse of our site's SSL certificate and favicon. Now about DNS. We have a domain name and we have the physical location of our site. Linking the first to the second is the job of DNS records. I mean, we all remember, some time ago, in 2021, a certain social network, banned and designated terrorist in Russia, suddenly stopped working. That was in October, I think. And journalists were writing to me, screaming, in tears: Igor, tell us, why exactly can you get onto this banned social network while nobody else can? I said I was going straight to the hosting, to the IP address of that social network. So that's what DNS records are responsible for.
Thus, DNS records hold all the information about the servers linked to the site we're examining. Some of them may not be covered by Cloudflare, some of them may be physically located on Russian territory. And we've come across cases where drug-trafficking sites, sites spreading false content, and other illegal resources may have part of their infrastructure based on the territory of our country. We came across quite a lot of that. SSL and favicon, which can also be used; listed here are products that can help us search for reuse of these site elements. And it's not quite a linear story when it comes to trying to identify and find the owners of our resource.
Linked contacts. We often look for contacts in the body of the site itself. And it's important to note here that contacts for a web resource turn up not only on the site itself; they can be in external leaks. First, registration data, archived data, WHOIS. It existed, it's been saved somewhere, it was in numerous scraping and parsing results, and it's available in large quantities on the internet. Second, there's a huge amount of advertising using one site or another, which can also be collected by various crawlers. And finally, the millions-strong leaks that happened here and worldwide; and those leaks can also be linked to domain names, so we'll see the pattern of how the email address is formed, which employee names appearing in those addresses exist on the company's domain, and maybe which passwords are used there. This lets us collect contacts.
On top of everything else, we can guess these contacts using a standard pattern. The domain name plus standard, for example, email addresses: office, contact, admin, support, HR, PR and others. Then check them with an SMTP request to see whether they actually exist.
Web resources also store a large number of external files. These files, of course, can be found using various external services. VirusTotal is your friend here. They can be found using advanced search operators, or dorks, like the ones on the screen right now. And most importantly, these files very often store metadata. One of our investigations, as funny as it may sound for 2025, involved an invasion of privacy. Two people, two businessmen, fell out. One made a website about the other and posted all kinds of scribbles there, nasty pictures featuring his competitor, and he made those pictures on an iPhone, and the iPhone dutifully saved in the metadata the geolocation of where those pictures were made. This was in 2025, as funny as that may sound, and the geolocation was in fact the home address of the person who was the main suspect.
Among other things, metadata, as we know, also stores data about the camera, the time and date of the shot, and other information that matters to us. Hyperlinks are another important part of the website under investigation. Hyperlinks can be both external and internal. Internal ones are links within the site, between individual pages of the website. And some of them may be hidden, not public. In that case we examine the robots.txt and sitemap.xml files to find out what other pages exist on the site that we can't see. They are often of interest to us. And finally, external hyperlinks. They can also be useful to us, because these are links to associated social networks, to external file-sharing services. For basically any file-sharing service, you can determine which email is tied to it, even by OSINT methods.
All this lets us identify the people who administer the site. For example, if we don't know and can't find the site owner's or administrator's details by simple means, but it has an associated group on the VKontakte social network, then using the utterly trivial InfoApp application we can obtain the profile data of the administrators of the group linked on VKontakte, and that will be much faster, more convenient and more probative. The next point, also quite important, is the technologies used on our web resource. What's included? Numerous bank acquiring services, chatbots, feedback forms, analytics counters, advertising ID codes, and so on and so forth. Everything we try to cram into our web resource in order to collect information about the audience, build feedback with it, in order to control it and build targeting — from a marketing standpoint all this works in our favor, but for reconnaissance on domain names and sites it's, of course, a huge minus.
At the very least, we've come across a great many Yandex.Metrica counters placed on sites — on banned sites, on sites spreading false information. And the funniest thing is that all of this can be used for identification. For example, we have a Yandex.Metrica code. First, about 10% of Yandex.Metrica counters are public: you can open one and simply look at the moment it was installed on the site. And when it was being installed on the site, naturally, the only user who came into its field of view was the person installing it on that site. And the second point concerning Yandex.Metrica is that through support — quite wonderfully — Yandex support often, not exactly as a secret, but discloses the email address associated with a given Yandex.Metrica identifier. Just ask them for a hint: "I'm the site administrator, I forgot which email is linked to this Yandex.Metrica counter," and they tell you that email address.
Things like that have happened too. Acquiring. Acquiring services are tied to banks. Then, with an acquiring service installed, you can contact the bank to find out who obtained that technology for placement on the site, and get his login and other registration information. Any site also exists in a certain historical retrospective, so we also turn to web archives to see what it looked like before, what the site's code looked like before, the elements and technologies that were part of that code, what contacts were previously listed on the site, linked social networks, external and internal hyperlinks, and even files. All this will be in the various web archives.
And finally, external traffic. It's often left out of website investigations, but for us the most important point arises here. For example, some sites hosting petitions. You've probably often come across petitions in your investigations calling to overthrow the government, topple presidents, governors and other officials. Here, very often, in about 70–80% of cases, we find that a petition, yes, can be inflated, a petition, yes, is almost always inflated by bots, but as a rule the petition's author initially tries to seed it himself. So we turn to the web resource's external traffic to find out who first posted links to that petition on the net with calls to vote for it. And we often find such people, find them on the social media page where they pushed it, the groups where they posted the calls with that petition, and that lets us find the authors of these very things.
To sum up this whole story, the overall outline of what we talked about today. Very briefly, because the full lecture on website investigation takes us about an hour and a half.
Everyone will get the slides. Useful sources, the bare minimum of what you can use within OSINT to check web resources.
And, going back, once again the link to the browser I showed at the very beginning. It has practically the quintessence of everything we talked about, both in our previous talks and in this one. Investigate cryptocurrency, social networks, websites, Telegram, work in a security department, do forensics — most of the software is generally free, and in this browser you'll find more than 2,000 sources.
Thank you, you've been a wonderful audience.
No doubt about that. So, colleagues, your questions. Igor, you've apparently fired up the audience so much that it's all perfectly clear now. Then let's send Igor off with another round of applause.
[applause]
And we'll move along, bit by bit, toward the very final part of our evening today, namely the roast.
Closing of the online part of day 1
Dmitry Yankovoy (moderator).
But before that, before we move on to it, now is the time to say goodbye to our online viewers. And we'll see you tomorrow, right at the start of our second day. And with you, dear guests, we'll now drift smoothly over to the bar area, taking along some drinks, spirited and otherwise.
[End of the day 1 stream: the "roast", the prize draw and the informal part were not recorded. What follows is the day 2 stream; it starts in the middle of the moderator's introduction to the first talk (with a pause in the recording between them).]