Day 1 — Thursday, 11 September 2025: digital forensics day
Conference opening
Dmitry Yankovoy (moderator), Olga Gutman (MKO Systems). Scheduled 10:00–10:10.
Moderator — Dmitry Yankovoy
Testing, testing. So, dear friends, hello everyone. Many of you know me, my name is still Dmitry Yankovoy, and today I'll be moderating our event. Introduced myself just in time. Two announcements down, high time for the third. So, last year, unfortunately, I didn't quite manage to make it here, but this year I'm before you again. Greetings to everyone who's watching us online, and offline too. I'm incredibly pleased to see a great many faces here, both old and new, especially — actually, I'll probably segment them, because the old faces I'm pleased to see for the reason that I see old friends I've known for many years, and the new faces I'm pleased to see because I realize that digital forensics as a topic is picking up more and more momentum, which means, I suppose, what we're going to talk about today is very important.
One important point — imagine, it said on the website, I don't know whether you noticed or not, that this is the ninth conference. We've been running this conference for 9 years now, since 2016, and next year, just imagine, will be an anniversary one. But in 9 years, only one thing hasn't changed, and it's the most important one. The most important thing in all of this isn't even what you hear from the stage, it's the connection that has formed between us. Because today, Moscow Forensics Day isn't just some conference on digital forensics, it's probably one of the few places where forensic experts from all over the country can meet up, talk and share experience. That said, again, I won't downplay the importance of the talks themselves. Ask questions. We've prepared, as it's fashionable to say now, some real meat, as far as the talks themselves are concerned.
But before I move on to the other housekeeping points, I suppose we should open the conference. So I invite to the stage our invariably wonderful Olga Gutman. Let's give her a round of applause, because without her, none of this would exist, 100 percent. Yes, you can pick up the mics over there. Olga, the floor is yours.
Opening — Olga Gutman
A very, very good morning to everyone. Both to those in the hall right now and to those glued to their monitor screens, possibly watching our conference in the middle of their working day. Like Dima, I'm very glad to see both familiar and new faces in the hall. Our once small, intimate event has grown into a large-scale conference. That makes us very happy. And even though we've turned into such a large venue where you can network, we always try to keep that atmosphere of coziness and intimate conversation. And most importantly, the usefulness of the conference, which has been the aim of this event from the very start. So throughout the day, be sure to talk to each other, don't be afraid to come up, don't be shy about approaching us. Ask the speakers your questions, including online. Or if you have questions for our company, or for the companies represented at the booths today, ask them online as well, they'll definitely be passed on to us, and you won't be left without an answer.
I'd like to thank all the partners taking part in our event today. Just a small aside. We started putting together this MFD's programme a whole year before the event. So in effect, today you're going to see the result of a year's worth of our work. And that's very flattering. We couldn't fit in all the presenters, all the speakers who wanted to present at our event today. But in the end, I think, what we've got is the cream of the crop. The talks aren't some theory that's divorced from reality, but topical talks aimed at practical application. I'm sure that today, together with you, and tomorrow, we'll have two extremely productive and useful days.
Once again, thank you all for finding the time to attend our event today. As always, I wish good luck to the speakers, partners and organizers of today's conference, and I declare Moscow Forensics Day 2025 open.
Moderator: the organizer, partners and the day's program
Olga, thank you so much. Well, and we carry on making today's history. As I said, I see a lot of faces, so let me start by introducing who we actually are. MKO Systems is a leading developer of software for computer forensic examination of mobile devices, personal computers, drones, cloud services and much more. We also run training courses in the field of digital forensics. All the details are in your bags. I'm not going to do any advertising now. But today, as you can see on the slide, we're by no means doing this alone. Alongside us today, you can see the booths of companies such as ACE Lab, LAN PROJECT, ELETEK, Ester Solutions, SearchInform and STC.
Also, our general partners for this conference, Account-Best, sadly have no booth, but I'm sure you can find Igor Evgenyevich in the hall, especially since he'll be giving a talk on UAVs today, and have a chat with him. Anyway, let's move on to some brief housekeeping. Just one request: I know you can't switch your phones off, so I'd just really ask you to switch them to silent mode. And if you happen to wander over to the booth area to chat during the talks, please don't be too loud. Actually, that's the housekeeping covered, and that's all of it.
Let's run through what's in store today. We've got three big, packed blocks ahead of us today, so that's three sections of talks, and then, as always, the hottest part of this conference: the roast. What is the roast? The roast is the session where you, as Olga already said, saved up questions all year, and now you can put those questions to pretty much everyone speaking today, that is, say, if you've built up something bad you'd like to say about Mobile Criminalist, that's welcome here, you can say it straight to the CEO, so it goes straight to the top.
Right, let me check. Yes, the housekeeping's all done, the conference is already open, which means we can move straight on to the talks.
1. Vladimir Greshnov (ELETEK) — “Duplicators vs. manual copying: when speed and accuracy are critical”
Scheduled 10:10–10:40.
Moderator's introduction
And our first section is opened by Vladimir Greshnov of ELETEK, with a topic that, I actually think, is similar to those, you know, even slightly philosophical questions that the famous American writer Isaac Asimov asked back in the last century, because, well, now that machines have actually, let's say, grown into the Skynet we saw back in the day in Terminator 2, it's become a pressing topic, and yet, who'd have thought, even ordinary duplicators are starting to nip at our heels. I hope Vladimir will clarify for us how things stand right now and whether we should be worried. Let's give him a round of applause.
Vladimir, you're on. This one here's for the next page. Here, take the mic. I'll go up here, I guess? Go ahead.
Talk and Q&A
Hello, colleagues. So, as Dmitry said, my name is Vladimir. I represent the company ELETEK. Today I'm going to tell you about one of the areas we work in. That's the design and manufacture of duplicators, which help us in our work with fast and safe acquisition of data from digital storage media without making any changes to them.
On the first slide you can see the products that will be covered today. It won't just be the hardware solution; I'll also talk about our software solution.
Many of you know our products. Some have already had them in trial operation, but some are seeing them for the first time. So the first system I'd like to tell you about is a system designed for copying data from personal computers and laptops without having to take them apart in order to pull out the internal drive, and without having to boot the native operating system, so that no changes whatsoever are made to that drive.
It's quite simple to operate. You connect it to the computer. Through the Boot Manager, you boot the operating system that's stored on the unit.
Once it's booted, the program mounts all the drives that are either installed inside the computer itself, or connected, say, flash drives plugged into the USB ports, and then you can proceed to acquire that data.
Now the slide will show the software part of this system specifically, what it actually looks like from the user's side once the unit has booted up.
As for new features, the most recent ones that have been added. Previously, we could only do file-by-file copying by masks, that is, we set certain masks, which you can see on the screen now, say, audio, video, or, more specifically, mp3, mp4 and so on, and when it scans the drive's file system, it only extracts those files, but now, in the latest release, we've also added search and data extraction by signatures. So if, say, an attacker took a document with a.docx extension and either removed the extension altogether or changed it to, say, mp3, or to something else, then the signature search we've implemented will help determine that it's actually a text document, and it will extract it as well for further processing.
Also, among the new features of this system, a lot of people asked us to implement parallel copying of drives. Previously, we had it set up so that copying ran sequentially: first the first disk is imaged, then the second, the third, but now you can select a drive that's inside the unit, whichever one you're interested in, and, say, either image it first, or basically image only that one.
Right, let's move on to the next product. The next product is our most compact duplicator. It was originally built at the request of field operatives so that during an operation they could, as discreetly and quickly as possible, image any flash drive that's, say, lying around somewhere: an officer could walk up at a convenient moment, plug it into our device, and it would automatically make either a sector-by-sector copy or extract all the files on that flash drive.
This unit runs on battery for up to two hours, specifically in copying mode. All of its pre-configuration, and the ability to see what's currently happening on it, whether copying is in progress, how many and which dumps are stored in the unit, is done through a mobile app. It's implemented for phones running the Android operating system. On the slide now you can see exactly what the screens of the app for controlling this device look like. But in principle, this device can also be used without a mobile phone. By default, it's set to create sector-by-sector copies of flash drives.
And the device itself has LED indicators that tell you the device is ready, that copying is in progress, and that copying is complete.
The next product is probably one of our newest, one of the most recent. It's a simpler type of duplicator than the next product I'll be presenting. This duplicator is designed to copy data from SATA drives. That's 2.5-inch HDDs and SSDs, or 3.5-inch HDDs.
Copying here is pass-through. To bring down the cost of the unit, we don't put in internal memory, an internal drive. But if needed, if the customer wants it, internal storage can also be added as an option.
Here, unlike on our largest workstation, there's a hardware write blocker, which is exactly what lets us avoid making any changes to the evidence drives we connect. During operation, you can see the device's status, the copying progress, and whether it's ready to go, on a small display. Like the previous system, this duplicator is controlled via a mobile phone. That is, selecting the operating mode, and you can also check the device's status. But it's also standalone, that is, when you connect an evidence drive, copying starts automatically.
Here on the slide are the same control app screens you saw for the previous product.
Our next product is our largest duplicator. It's a full workstation with its own large 7-inch touchscreen for control. This system is built on a Mini ATX board with a full-fledged processor. In this particular version, the one in the photo, it's an Intel Core i3. It comes with 8 gigabytes of RAM installed. If needed, that can be expanded to 16 or 32.
For internal storage, there's a 2-terabyte NVMe drive, which can also be upgraded to 4 terabytes if needed, or, basically, whatever the NVMe drives themselves allow.
This unit, the reason we call it a full-fledged workstation, is that it can be used both in the field, for example, out on site, it can run off a power bank, and also in a lab setting. In terms of operation, it lets you copy from external drives in sector-by-sector mode, in cloning mode, that is, a full disk-to-disk copy, in file container mode, and you can also pick, through the file manager, specific files you're interested in, and it will extract only those.
Of the new things we've done: I think at the last exhibition, or the one before, we were asked about the size of the block it uses for sector-by-sector copying. Here we've given the user a choice: you can select a block size from a minimum of 64 kilobytes up to a maximum of 8 megabytes.
We've also reworked the hash calculation feature a bit, made it more, so to speak, complete. Previously we only did a fixed verification, but now, accordingly, if the user needs it, we can either calculate hashes separately, or calculate hashes directly during the copying, the duplication of the evidence drive, or also run a verification afterwards.
We've automated some functions that were previously available only to a user who was, so to speak, trained, who knows that he needs to, say, check a disk for hidden areas. Now all of that happens automatically. When a copy job is added, the unit warns the user that hidden areas are present. To disable them, the user has to go into the settings. And we've also added a user notification about disk encryption. For now, we've implemented a BitLocker notification specifically: if a disk is BitLocker-encrypted, you can only take a sector-by-sector copy of it, but we're also currently working on other types of encryption.
On this slide you can now see what the visual part of the system's software and its controls look like.
For convenience in working with the unit, and later on with the copied images, we've also implemented a report generation feature, where the user of this device can look and, so to speak, refresh their memory as to what mode the copy was made in, from which disk to which, whether everything completed successfully, what the hash is, if hashing was enabled, and so on, everything about the unit's operation.
Here on the slide you can see a status drawer that shows the mode in which these drives are connected. As I said earlier, evidence drives connected to the unit are connected in read-only mode, which is indicated by the closed padlock in the status drawer. If needed, you can also use your own trusted drives, to which data will be written, or exported from the internal storage.
For them to show up, we've implemented a special volume label, when it's set, the unit understands that this is a drive it's allowed to write to, and it doesn't block it for writing.
Various disk operations are implemented here, such as restoring a previously created image to a drive, or rather, deploying it, if the user needs that.
Let's move on to the next one. Here, as I said at the beginning, it's not only our hardware solutions that are presented. This slide shows our software solution. It's a software package built for viewing images in various formats, such as RAW, our most standard format, and E01. With the ability to recover deleted files, view deleted partitions, and detect deleted operating systems.
In the latest releases of this product, our main focus has been on Astra Linux, so that it can be used on that operating system. And most importantly, we had a request for viewing the E01 format on Astra Linux. That's now been implemented as well.
This slide shows the visual side of the program. Basically, all the systems, all the programs you're seeing in my talk today, you can come by our booth and we'll show you how they work in more detail, or tell you more about each of them.
And our last product for today is also a software product. In how it works, it's similar to our very first product, which is designed for extracting data from personal computers, but with one key, main difference: it works on a live system. So, if needed, you have a computer that's switched on. With this software package you can go through the computer's entire file system either in mask mode or in signature mode, and extract only those files, only the information that you actually need. This package also runs on Windows and Linux systems.
For Astra Linux, that's still in our plans, but it's coming too.
Here's what the control panel of this package actually looks like.
But again, as I said earlier, you can come by the booth and we'll show and tell you more. That basically wraps up my talk. Your questions, please. So, colleagues, let's do it by a show of hands. Ilya, I'll come over to you. —
— Good afternoon. You didn't mention, what about RAID? —
— Look, as far as RAID goes. You can image RAID as well using our first system, which is called Element-U. But after that, it's up to you how to reassemble those RAIDs. So the whole problem will be in the subsequent reassembly. Yuri Mikhailovich, go ahead. Good afternoon. So, about the first unit, the one that connects directly to a computer. Have you done anything about RAM dumps? No, we don't work with RAM yet. That's bad. I mean, we're, so to speak, on the verge of starting to work on that, but the main focus was on working with file systems. RAM is something we do want to add.
Next, one more quick one. So, that is, drives that are under BitLocker or something else. Will you do anything about those? Say that again? Encrypted, encrypted. What do you do with a clone of the drive? I see what you mean. Look, as far as encrypted drives go, you can make a sector-by-sector copy of any encrypted drive. But actually viewing the data with our devices — that won't work here. And one very last question. Why make a clone at all, if an image is better? —
— By clone you mean disk-to-disk? Yes, yes, yes — an image, I mean. First of all, we can compute a hash. Here, hashing is pointless for you. And then do whatever you want with it. Have you implemented imaging at all? Look, actually our main mode, the one we started the development of every unit from, is a sector-by-sector copy to a file. We added cloning here just as an extra, because a lot of officers asked to be able to make a copy and immediately examine the information that's on the drive. Here it's more of an add-on. —
— Colleagues. Ilya, hand over the mic on that side, please.
I'd like to clarify something about the Element-U. Could you bring the mic a bit closer? About the Element-U. You said we connect it automatically and make a dump. A bootable operating system is loaded. What if Secure Boot is enabled and a TPM module is in use? I mean, does it boot, does the system get dumped? Look, we do have a Secure Boot bypass in this software specifically. But as for the TPM module specifically, it's hard to say off the top of my head. You can come by our booth. We have the engineer there, the actual developer. So you'll be able to discuss it with him directly. Colleagues, are there any more questions? Ilya, over there, in the front row. Thank you very much. Hello, tell me, if, as you said, for operatives, several USB drives onto internal storage, is it possible to save them all at once as images? Well, look, as for — I take it you mean the device that acquires from flash drives. Yes, yes.
It's designed to acquire only one storage device at a time. Thank you. —
— Right, colleagues, any more questions? No more questions, then — I see no hands. Vladimir, thank you very much. You can leave the clicker and the mic right on the lectern. Let's see Vladimir off with a round of applause. That was great. So we do still have a chance after all.
2. Valeria Vakhrushina (MKO Systems) — “Dictionary or mask: how MK Brute Force works”
Scheduled 10:45–11:15.
Moderator's introduction
So, before I introduce our next speaker, let me clear up one thing right away: there is no magic today, no sorcery, no wizardry. Maslenitsa is over, and we won't need to burn anyone today. Let's dot all the i's right away.
It's just that, you know, the password "password" still comes up quite often. But all of that will be covered today by not just our marketing director, but the head of Brute Force module development for Mobile Criminalist, Valeria Vakhrushina. Let's give her a round of applause.
Dear beloved boss, your microphone and your clicker.
Talk and Q&A
He made me come up on stage. Good afternoon, colleagues. I'm very glad to welcome you. Yuri Mikhailovich, please don't bury me in questions today. Thank you.
So, today I'm going to tell you about MK Brute Force — dictionary or mask, how it works. Dima's already introduced me; you all know my split personality as a marketer, but I also work on this module. We live in an age not only of digital forensics, but also of total digitalization. What do we have? We have loads of apps, phones, tablets, smartphones, laptops, and we put passwords on all of them. That's as safe as it gets, but I also know that both security folks and IT folks always say: please, set proper passwords, please make them long, use mixed case, use special characters, letters, make passwords at least 10 characters.
Well, I understand that among you, yes, surely everyone follows that advice. Unfortunately, among ordinary users — and fortunately for us forensics people — it usually isn't followed. Here are the most popular passwords. Well, we all know qwerty, password, guest, admin, admin — nothing's off the table. How often do you think they're used? I found some very curious statistics that horrify me. But then again, for brute-forcing that's good.
And then there are people who think they'll use some simple combinations. Like, a digit, two letters, the next digit, the two adjacent letters. Straight along the keyboard — very convenient to type all that. We understand perfectly well that if we run a mask attack, and we understand what the combination is and how to define it, then that password will be cracked just as easily. I also want to show you some interesting statistics on the most used passwords in the Russian Federation. The stats aren't the freshest: 2023–2024. I especially liked "Baltika 9" and "Sotochka". Really great passwords. But home phone numbers, cell phone numbers — of course, again, a lot of people use those. Why memorize something extra?
So, the usual advice? Change it regularly, the password must consist of at least refrain from using, blah-blah-blah-blah-blah. I recommend to everyone — to you, not to those whose passwords we'll be cracking — I recommend to everyone: 2FA, always. Okay, let's talk about password cracking itself. Whatever we're cracking, whatever we're cracking the password to, in any case, in any tool, it always comes down to two methods. It's either a dictionary or a mask. A dictionary, in principle, also counts as brute force, but how does it work? We have some dictionary pre-loaded into the program — it could be our MK Brute Force or some other solution, open source, paid, doesn't matter.
And it contains a certain set of characters, a certain set of strings. In our case, for example, it's a txt file with the most popular or, say, leaked passwords, or ones you created yourself, and so on. And the attack runs directly through those combinations that are defined. What matters? A dictionary is a great thing, of course; an attack with it will go much faster, but if the password is, let's say, safe, following all the recommendations of IT and security people, we're unlikely to crack it with a dictionary. That said, I want to note that it's a must when cracking passwords to phones. Because there we're dealing with either PINs or patterns, and in that case a mask takes longer; easier to run it all through a dictionary.
I know there are lots of tools now — again, that's what they write online — that people have started using AI to create new dictionaries based on some leaked data. Well, let's say, applying some mutations in advance in order to later load that into some brute-forcing tool and crack the password. A cool thing. We made something roughly similar. In Mobile Criminalist, in MK Brute Force, we call it a dictionary based on personal data. What is it? If you choose this method, a password manager opens, and it holds as much data as could be gathered from the extraction. That's personal data: the first part of emails, phone numbers, surnames, etc. Based on that, you can build a dictionary, which, again, can speed up cracking. But here you do have to do some magic.
If we're talking about something complex — the mask. A head-on mask — I'll say it right away, don't use a mask head-on. That's 26 letters of the Latin alphabet in upper case, the same number of letters in lower case, plus 10 digits, plus 33, I think, special characters. So if you launch such an attack head-on, we need either super-powerful hardware, or, I don't know, or there has to be some element of luck. Just as an example. A ten-character password made up of just lowercase letters. 141 trillion possible combinations, well, more than that.
Now we add uppercase to that. At this point I even — sorry — googled what these numbers are called. More than three quadrillion. Let's add digits. At this point we're already talking more than a quintillion passwords. Well, sorry, I'm used to working in powers, so I'm checking my notes. Well, basically, cracking such a thing, I don't know, is unlikely to work out. That's why we have standard commands. I specifically put the commands up — how to create a mask, how to set it. Makes things a lot easier. Basically, these are hashcat commands, applicable to MK Brute Force. So in this case, name is a fixed value, either at the beginning or at the end.
Next we've got a question mark plus an a — that's basically any character. So there's room to play around here. Well, I'll show you right now on video, using MK Brute Force as an example. So, we know that if we need to crack a password for something as part of some investigation, we're usually very tight on time. So for my part, I'd recommend: first we run through a dictionary, then we add a mask. That'll make life easier and speed things up a bit. And a second point. GPU power, CPU power. We've got hashcat built into MK Brute Force. So our password cracking runs faster on the graphics card. In some tools, password cracking runs faster on the CPU.
And of course, if your hardware allows it, distributed password cracking is a really great thing. It'll speed up the password-cracking process a lot, and you'll get a successful result much sooner. Anticipating your questions — we're working on it, since we've had lots of requests from you, but we haven't implemented it yet. A lot of tools have it. To sum up: cracking isn't just brute force for its own sake, to mess around, you really do need some meaningful hypotheses, to use some meaningful dictionaries, and keep all of it updated in time, so you don't sit there running hardware and burning electricity for nothing.
I'll show you a bit on video. So, what I was saying about the combination. Here I'm taking a standard ZIP archive. The hash will be recognized automatically, so with MK, we just load it in there. It takes some time to extract.
In the first case, we go by dictionary. I take the simplest preloaded dictionary — the 100,000 most popular passwords, no mutations, just head-on on the GPU.
With ZIP, we'll see the result during initialization — we won't even see the attack process, because ZIP is a very simple hash. So besides what I mentioned — the cracking complexity, the mask complexity, GPU power, the amount of GPU memory — what else matters? What also matters is the complexity, the heaviness, say, of the hash we crack. Because a ZIP archive will always crack fast. If we're talking Telegram Local Passcode or some Huawei HiSuite, things like that — of course, those use completely different encryption methods. My colleagues will talk about this a bit later — that'll be the next talk.
There, cracking will take much longer. So, we can see the dictionary run has finished. Not a single password matched. Unfortunately, no success. And so we try a mask. Here, of course, I know this ZIP archive's password, since I recorded the video. So I know right away that I've got 7 characters there. And that's how I set them up. And I suggest running it again. I know the password has caps and digits. Now let's just try it head-on. We're not going to crack it right now. We won't wait for it, I'll say right away. You'll just see the amount of time you could potentially spend on a task like this.
7.5 hours. Sounds scary, right? Now let's try a mask that's partly custom-written. Again, the commands are standard hashcat ones. Those who haven't tried using this — give it a try. It's a really great thing, it makes the task a lot easier. So at the start we've set 3 completely random characters, and then 4 fixed digits.
Just in case anyone has questions later, why initialization takes a while, if you've used the brute-force tool — and I hope you do use it — again, it's due to the heavy hash. So you can see, during initialization we got the password MFD2025.
We'll see exactly the same picture if we do it the other way around, that is, write a mask but specify the first three letters, then leave four digits to be cracked. It'll crack just as fast.
But for the full picture, right now we'll wait here for it to initialize, we'll make sure MK Brute Force works, and I'll be pleased.
There, we can see the password was again cracked in 11 seconds. Well, pretty cool overall, considering I can't say I've got the fanciest graphics card.
But for our real-world conditions, we'll again set it for 7 characters. This is what I'm creating right now. But the first three are completely random, and the last four are digits. And now, during initialization, we'll already see that it's not 7.5 hours, like when we run a mask over everything possible and impossible. It's, I think, something around an hour. Yeah, there — 51 minutes, a pretty decent hashrate. But in that case it's already worth a try. There's a chance of success. Okay, so, regarding MK Brute Force. I hope you've all used it. Let me remind you it's built on hashcat, which is considered one of the fastest solutions on the market.
I know hashcat 7.0 has come out — we'll soon build it into MK Brute Force. It's got quite a lot of updates that'll interest you. The main goal, actually, was an easy interface, to lower the barrier to entry. Because hashcat, if you're familiar with it, doesn't really have an interface. This is what MK Brute Force can currently crack passwords for. So, the super-useful stuff, I'd probably highlight this for you, I suppose — it's Androids, iTunes, HiSuites, and other cool things.
People often ask questions, so I put it on the slides, in the presentation. You need to know the hash in hashcat format. Where do I get the hash? For archives and office documents, we load them right into MK Brute Force, the hash is recognized. For a passcode for Telegram Desktop, BitLocker, NTLM — through Scout, Scout passes the hash to MK Brute Force automatically, and you launch the attack. For all the rest — that is, all mobile device backups plus Apple Notes, you load the backup into Mobile Criminalist, and a modal window will pop up, asking you to enter the password or crack it. When you click "Crack", again the hash is recognized, and you launch the attack.
Those who haven't used it and haven't bought MK yet for one reason or another, we have a free mobile version of MK Brute Force via the QR code. You can use it — it's specially for you. Oh, so many phones — how nice. Wait, wait, wait. Let me say right away that, if anything, you can go to the website, and there, in a separate tab, there's MK Brute Force, which you can download from a computer. Absolutely right.
Alright, wonderful. Use it. I'll be glad to hear your feedback. I hope there'll be a minimum of bugs. Thank you for your attention. I'm ready to hear your questions, wishes, suggestions, what's missing, what you'd like, which hashes you want supported. And if anyone wants a closer look at something, you're welcome at our booth, let's talk. The suggestions, I suggest we save for a bit later, for the roast, and move to questions for now. But before we get to them, I really loved your line when you said you don't have the best graphics card, and you've got five there. That was really good. I really don't have the best graphics card, some of us have the Ti ones.
Alright. Colleagues, any questions? There are. Ilya, go over there, please.
Two quick questions. Is the functionality available in the Brute Force application from the link? Specifically, is password cracking of a physical Android image available? And the second question. If automatic password cracking doesn't start on that same Android, we import the image into Mobile Criminalist, but it doesn't start automatically. Can it be set up for brute force manually? —
— FBE and FDE, that's my pain point, sorry. No, unfortunately, FBE and FDE aren't supported in the free version. I can tell you right away that there are specifics to hash support. We built in the most current versions of exactly these image types. So if something specific isn't supported and the attack doesn't start from Mobile Criminalist in automatic mode, then write to support, making sure to specify the device, because I'll need to do further research and add support for those hashes.
It will be possible, of course. We're happy to add to the backlog. Ilya, tell me please, is that all over there? Next question? —
— Wait, Ilya, Ilya, Ilya, one second. San Sanych is there. Ilya, one second, I also have a very... —
— Good afternoon, is distributed cracking ever going to happen? I did say, we're working on it. You're asking really hard, we're working really hard, honestly. I'm hoping for the first half of '26, but I'm not promising anything. The dark circles under your eyes confirm it, right? No, honestly, we started working on distributed cracking, but since hashcat released version 7.0 with a huge number of updates, which will be useful to you, so we decided to update the core first, after all, and only then deal with distributed cracking. Our developers have already lost 5 kilos. Ilya, there was another question over there. Valeria, I listened carefully. You were saying, on the one hand, there are possibilities for AI to create non-standard passwords, right?
Have you modeled the situation the other way around? AI capabilities for figuring out passwords, your developers in particular? —
— Honestly, no. Unfortunately, building artificial intelligence into Mobile Criminalist is prohibited. The legislation would object. But what's the difference? You'll get there anyway. No, San Sanych, honestly, not yet. The idea is interesting, trying it the other way around, but this is probably closer to rainbow tables. Anyway, we'll think about it. I see a raised hand, I'm on my way. Only thing is, I'll probably go around you from this side, because I might not squeeze through over there, so I'll need a little help. —
— Hello, a question has come up. You just said that the legislation prohibits the use of artificial intelligence. Can you say which specific provision prohibits it? Because from what's been said, it's not quite clear what's meant by the artificial intelligence used to crack passwords. And there's a chance that actually what's meant by that is just some traditional algorithms, the same ones used everywhere. And it's unclear how, in that sense, the legislation comes down with total clarity and says, oh no, that's not allowed now. So, I hope the question is clear. —
— The question is clear. Of course, I do have a law degree, but I won't answer this question from a legal standpoint. But I can say that Mobile Criminalist by itself doesn't analyze or accumulate your information. We hand you the software, and we get nothing back. So, to train anything, we'd need to get information back from you. And that would not be good. —
— Now it's clear.
Dmitry Yankovoy, are you going to let me go today? One more question. Alright, while Ilya Anatolyevich gets to the question, I have just one request, colleagues: all of you, literally all, will get a recording of today's event, so please put your phones away and don't film. We'll email everything to you later, it'll all look nice. Ilya Anatolyevich, the question now. Just one more, really quick. You mentioned that you can build a dictionary based on the data extracted from a specific device. In practice I haven't seen how that can be done from Mobile Criminalist? Yes, it's from Mobile Criminalist, that is, when you're in MK Brute Force. There are three methods. Standard dictionary, mask, and a dictionary based on personal data. If you select the dictionary based on personal data and click the "Add" button, it brings up exactly the password manager built into MK, and there you can build that dictionary.
Okay, I suggest we leave everything else for the roast. Valeria, thank you very much. Let's send Lera off with applause. We're leaving everything exactly as it is.
3. Vyacheslav Chikin (ACE Lab) — “Specifics of accelerating password brute-forcing on mobile devices”
Scheduled 11:20–11:55.
Moderator's introduction
Well, I think password cracking has become more or less clear to us all. But now let's narrow down the topic we're talking about a little. Because overall we'll talk about mobile forensics, which is what we started with. Let's narrow the topic. There are always lots of theories that our phone is basically a safe that stores information. And the best safe there could possibly be. But let's remember, a safe generally doesn't run out of charge on its own. And sometimes the battery drains much faster than you could ever crack the password on that phone. How to speed all this up and how to do it all faster, our next speaker, Vyacheslav Chikin from ACE Lab, will tell us. Let's welcome him with a round of applause.
[applause]
Vyacheslav, by and large, all our hopes rest on you. Here you go, your microphone, the clicker.
Talk and Q&A
Good afternoon, thanks, Dmitry. Let's continue the topic of passwords.
The way things are going, device power is growing, capabilities are growing. Whereas before, 70%, or maybe even more, of cases were pattern locks. At most, a short PIN. Now, with the advent of face recognition, fingerprints, and maybe they'll come up with something else, it sometimes happens that even the phone's users themselves forget their password. And often they come up with a lot. Very long, complex passwords, a standard situation.
The person hasn't typed the password, and then the phone asks them to enter it.
This is especially common among teenagers. Recently my son came to me and said: Dad, I forgot my password. Can you crack it? An eight-character password, I think, no big deal, let's give it a try.
I hooked up our system, the phone is supported by our system. In the end I struggled a whole day and still haven't cracked it. I had to dig deeper into this topic.
So, first let's figure out what we can use as password characters. We have digits, the simplest case. Lowercase English letters, uppercase English letters, and special characters. 33 of those, as already mentioned. In total, that's 95 possible characters. But that's not all.
On some phones you can enter various special symbols in a password. You can enter a period, accented letters, little hearts and all the rest. Let's just say, Gen Z are actively using this now. And accordingly, this has to be considered.
But it's not all that complicated and sad here. Because if we take that same period and use it as one of the characters, then during brute-forcing it gets replaced with quotes like these. The heart symbol will be replaced with the letter "E". You can enter either the heart or the letter "E". Keep that in mind. But at the same time, this changes how you build masks for password brute-forcing.
For example, someone might set their password to "Ivan loves Dasha".
On mobile devices, the most common algorithms in use are SHA-256 and scrypt. As for SHA-256, everything's clear there — there are also, let's say, ASICs that you can get to crack passwords, to speed things up. As for scrypt, we'll talk about that a bit more in detail today, because it's the main algorithm used in mobile phones.
So, what is, let's say, the main difficulty, or even, let's say, the nasty thing about scrypt — it's that this algorithm is very hard to speed up and parallelize, because it works with memory. That is, initially some block of data and some salt, generated completely at random, goes into — let's call it that for now — a mutation block, and from it a big, big block of data is assembled; then from that block of data, at random, per the algorithm, blocks are picked, from them a hash is computed a certain way, and in the end we get our scrypt hash, which we brute-force the password against.
So here we immediately see two problems. The first is that we use memory heavily. For example, a graphics card may have loads and loads — thousands of compute cores, but we'll hit the fact that we eat up all the graphics card's memory, and those cores will be jostling, fighting over memory, and we won't get any speedup, any efficiency at all. The second bottleneck of this algorithm is, again, memory. It's that all this data, these random memory requests, start piling up in the memory controller. And we may have a super-powerful modern CPU in there, but all that data just hangs at the bottleneck, the memory controller.
So, let's talk a bit about the scrypt parameters. The first parameter is N, the so-called cost factor. It's the main scrypt parameter. It's all written here, I won't go over it. You can take a photo of all this. The slides will be available afterwards, you'll be able to look through it all. For now, we need to understand that this parameter — a great deal depends on it. That is, the higher this number, the more memory you need, the harder it is for the CPU to brute-force.
It's already noted here as a peculiarity that doubling the value of N doesn't just double the running time, it increases it roughly fourfold. Then r is the block size; that, actually, can be varied if we want more memory or less; in practice it's usually 8.
p — we'll skip that for now; it's for when you need to parallelize and somehow account for parallelizing the algorithm.
So, here are some quick calculations, to make it clear. Now let me explain. This is the formula, all boring, I'll put it a bit differently. So, there's this whole thing now called ASICs. And we too, when we first got into this, thought: wow, that'd be great — forget graphics cards, CPUs, here are ASICs, cryptocurrencies, it's all the rage now, let's hook one up, hack an ASIC somehow, bolt it on, make it work for us. But it's not that simple.
ASICs run — well, take a coin like Dogecoin, it and the others, basically, they all run the same scrypt algorithm with parameters 1024, 1, 1.
But that takes only 130 kilobytes of memory. So everything mines just fine, coins get minted, but for our tasks, unfortunately, that doesn't work.
Our task is a bit more interesting. Our parameters are 2048, 8, 1.
That's for Android with file-based encryption. That's on modern phones, and here a single CPU core or a single GPU core will already be consuming 2 megabytes of memory. Plus, those 2 megabytes, with random access, will start clogging the memory controller. So all of this has to be taken into account. For full-disk encryption — that's for older phones — we already need 32 megabytes of memory per core. If you've used our system, you've seen that on older phones the password takes much longer to crack. That's the reason.
Our system, too, is gradually starting to support parallelization. We ran it on some small setups, made builds, tried cracking passwords.
And we got different results for Windows and for Linux. Possibly because memory is organized differently. We'll have to look at how to optimize that. But on Linux we got much higher numbers — double — on the CPU. As for graphics cards, the results there are the same. If you look at the bottom row — that is, we've got, we tested a Ryzen 9 9950X, and a Core Ultra. Basically, despite their different number of cores, the results came out the same. We're still going to look into this, but most likely, our assumption is that we've simply hit the memory controller wall. It's got DDR5 in there. We'll keep experimenting going forward. For now, these are the results.
The graphics card, for those interested — a 4060 Ti showed only 1,500 passwords. So the logical question here is: what's better to use, graphics cards or CPUs? Well, let's say, whatever you've got. You can basically brute-force on anything. Down the line we'll have the option — say, if there are 10 computers in the office, all of them can be put to work on a single job.
So, back to it. Based on what we've learned today, let's try to picture what we can expect. Say, even if we build a more or less decent machine, it'll do 20 thousand passwords per second. That'd be either two good CPUs or a whole stack of graphics cards. To exhaust an eight-character password, we'd need a full 10 thousand years. Some people count in exponents; we prefer to count in years. Or in millennia, for now.
If there are any questions, I'm ready to answer them. —
— Thank you very much. Right, colleagues — ah, I see a microphone right away. Oh, a hand, I mean. Mic's with me. —
— Good afternoon, thanks for the talk. I have a question about memory — do you mean the CPU cache or RAM? Which memory is being used? —
— We're still going to be researching that, because, to make it… Right now I'm talking only about the CPU, not the GPU. To make it work with the cache, there are some options there too, certain tasks. That is, you need to write specific code in assembly.
We haven't gone down that deep yet. We use standard functions. How they get spread out there, we don't check yet. But yes, the option of using the cache — such options do exist. But then again, the cache isn't that big. So even if it's 2 megabytes, and with 24 cores on modern CPUs, there's a good chance it gets eaten up very fast. Yes, some threads could be sent there; possibly that's what's happening already. Well, let's say, that's a separate research topic, of course. Yes, good question.
Well, it's not really a question, more of a wish. You're known, you've always been known for your ability to work directly with memory, devices, processors. But something like a master password, or the like — haven't you tried digging in that direction? Like in Windows: take the password, change it — do that on a phone. Haven't tried? And aren't going to? I didn't quite follow. Well, look, we change the password. I've got a user, I change the password to my own and then operate with my own password. Have you tried it that way — you can't? So not guessing the password, not brute-forcing, but setting your own. Setting my own password? Yes, but keeping his data. —
— Well, setting your own password is hard. No, but it's not 10,500 years, is it? Maybe it'd even be easier? Ah, well, I was just looking at it from the hardware point of view, and the colleague before me explained how to cut that time down. But for that, again, dictionaries exist, and beyond that it's a creative process. —
— Bypassing it? Well, bypassing is hard — it's math. The math here — take SHA-256, it's a very simple algorithm too. But so far nobody's learned to find collisions for it. There are miners, whole farms, whole cities built to hunt for collisions.
Right, colleagues, more questions? If I see no more hands, let's send Vyacheslav off with applause. Vyacheslav, thanks a lot. You can leave it all over there at the booth. And we've all done great, because we've now reached the end of the first block of the first day of our event. And now we'll have a fairly long break, because, again, as I've already said, the most important thing is networking. You'll have a full 45 minutes to talk. The catering out there's basically laid out. And at 12:05 we'll meet back here in this hall. Thank you.
[A break (12:00–12:30 in the program) is cut from the recording; timecodes run without a gap.]
4. Olga Tushkanova (Main Forensic Directorate of the Investigative Committee of Russia, GUK SK) — “A standard methodology for examining information stored on mobile devices and their components”
Scheduled 12:30–13:00.
Moderator's introduction
We're moving on to the second part of today's event. Once again, I invite everyone to come back into the hall. You can do that with your coffee and sandwiches. Personally, I don't see any problem with that. As we've already gathered, broadly, from our previous session, the phone, even the one in my hands right now, really is a genuine black box, since it stores everything you searched for, who you called, what you browsed at night, and so on. And overall, even once we've unlocked this phone, we need to understand what's in its contents. Olga Vladislavovna Tushkanova will now kindly tell us about the standard methodology for examining precisely this kind of evidence that arises in these cases. Let's give her a round of applause.
Olga Vladislavovna, the mic is yours. I'll bring the clicker over now.
Talk and Q&A
Good afternoon. Oh, well then, yes, thank you very much. Right, where's the... the up arrow, that's that way. We'll sort it out now. Why did this topic come up? Actually, all of us, as forensic experts, are, generally speaking, obliged to cite some literature, some sources that we rely on when conducting a forensic examination. This is an integral part of the expert's report, where it must be stated. One, two, three.
No.
So, I'll now take two sources of this information, where an expert can draw from, and what he can refer to, take information from for his work. In fact, there are scientific publications, where all sorts of things get written, and non-scientific publications, various recommendations that developers produce, and all the help files and the rest — it all exists. But the main thing experts use is either methodologies, or methodological guidelines.
A methodology is a formalized algorithm of the expert's actions, ensuring the reproducibility and reliability of the results. It includes the stages of preparation, execution, and evaluation of results, as well as requirements for the tools and the conditions of the examination. So that's what we call a methodology, and it's the main thing an expert should rely on. Two years ago I talked, more or less, about which tasks absolutely require methodologies, which tasks they're optional for, and which tasks are, well, where a person just writes: it's a search task, I can search this way or that, here's what I found, there you go, or didn't find it, so be it — like a crime scene examination: whatever you found goes into the criminal investigation pool; what you didn't find, you didn't find.
Methodological guidelines are advisory documents that explain and help apply forensic methodologies and equipment. If a methodology is one, two, three, five, mandatory, then methodological guidelines as a rule, explain how to apply that methodology. For example, our standard methodology for examining computer information, already written and improved several times. And under this methodology for examining computer information, we've now started writing methodological guidelines. One of them: methodological guidelines on examining information in Astra Linux. Astra Linux: it details the trace picture; the methodological guidelines tell you where to look, what to look with, and how to look, more specifically for Linux. Now we're writing the same for macOS.
But unlike that, we started creating methodologies, I suppose, more by object, so the methodological guidelines on examining computer information, they mostly concern information that resides on regular computers, in the form of file systems. It's a different matter when it comes to methodological guidelines and methodologies for examining mobile phones. Let's start with a bit of historical background. Why the three dots?
I can't claim to know everything that's been written on this subject since mobile phones first appeared — who wrote what and how. Including developers of certain software — they also described some things. Do this, do this, do this. The first more or less reliable source is the methodological guidelines on the forensic examination of cellular mobile phones in the bodies for control over the trafficking of narcotic drugs and psychotropic substances. I had the 2011 version; the FSKN — the drug control service — developed it, and these methodological guidelines were marked "for official use only". So they reached some people and didn't reach others. The drug control service used it itself. As far as I understand, one of the authors of that document is actually here today.
The next thing we got — guidelines, guidelines, we needed a standard methodology for examining information. And in 2014, the EKC MVD of Russia more or less wrote such a methodology. But it was written to the level of the objects that existed back then, and many of the things that have since appeared in mobile phones, this methodology really couldn't take into account, because it was hard to imagine what else would show up. The next thing, done in 2023: the EKC MVD of Russia improved this methodology a little, but it doesn't, let's say, differ much from the last one. And this year we wrote a standard methodology, developed it, for the forensic examination of information contained in mobile devices and their components.
I won't list for you one, two, three items and lay out how it's written. I'll just tell you about the innovations that were made, and from that it'll become clear why, generally speaking, we took this on. Why did we have to write a methodology for this kind of examination?
What's new? First. Previously it was just mobile phones. We've moved to mobile devices. So we've broadened the range of objects of examination a bit, because the principles and approaches are the same. What else did we include? Keypad mobile phones, obviously, smartphones, tablet computers and smartwatches, smart bands. So we've broadened the range of what falls under this methodology, right here. We reworked the section dedicated to the objects of examination, that is, their description. A standard methodology necessarily consists of several mandatory sections, including a brief description of the object of examination, the standard questions, the expert task, approaches to the equipment you use. The methodology itself — do one, two, three, four, five — but without being tied to specific equipment, just recommendations, so you don't forget to check here, check there, and check over there.
And the conclusions and so on. We've also added a section there on describing the objects of examination, describing technical data encryption. Among other things — much is new there.
But the requirements for the equipment used to conduct the examination have been clarified. So, let's say, the problem of what we should record in the methodology for examining computer information specifically, how much detail we should give about the equipment. If we go down that road — say, today I'm examining a computer or a mobile phone, and I need, preferably, this specific bench equipment, these specific hardware tools. Preferably, I should have something from ACE Lab, from MKO Systems, from someone else. Plus such-and-such software by name — then that's a road to nowhere, because in six months the names of the software products will change, something new will appear that I want, and the thing I wanted will disappear, but, unfortunately, they've simply gone off the market and aren't sold anymore. So the requirements in methodologies are written as functional ones. And what they'll be implemented with at any given moment, what you'll order and put into the technical specifications for procurement, into the contract, that will be something that meets these requirements.
But it will take on a concrete form and a concrete name of software and hardware. Equipment requirements: the expert's hardware and software workstation. Look: have various interface and technology connectors, specialized devices — well, a Bluetooth adapter, a Wi-Fi adapter, some other equipment — interface cables for connecting mobile devices, SIM cards, memory cards. What exactly, under which name right now — I have no idea, but it has to be able to do this. Have equipment for working with information on digital storage media. Have the ability to block the registration of mobile devices on cellular networks.
Have the ability to provide electrical power to batteries and directly to the mobile devices themselves. For example, a universal power supply unit. For example — go look, maybe something else will come along. Have the ability to remove memory chips from mobile devices. Why, how, which product? I won't say. The capability has to be there. Have the ability to read information from the memory chips of mobile devices. Next. Have the ability to read SIM cards and to write SIM card memory identifiers. Have the ability to decode the user partitions of mobile device memory. Have SIM cards that are not registered by a cellular operator.
Have an ultrasonic bath and drying equipment. It says "optional" here, meaning there are cases when a phone has to be properly washed and dried. Have the ability to search and interpret information obtained from the memory of mobile devices, memory cards and SIM cards. Again, here they tell you what they interpret it with, how they crack passwords, doesn't matter. The expert has to have that capability. Have the ability to video-record actions performed by the expert during the examination. That's sometimes needed, and among other things it must illustrate certain points. And have the ability to document the examination results and to prepare report files. Nowadays practically the majority of software products we have for examining information on mobile phones generate report files. So, well, let that capability be there.
The next of what's new for us. Well, let's say the examination procedure has been made more specific and expanded. Among other things — this isn't all of it, there's actually a lot, no point listing it, it's not interesting — we've added measures ensuring the preservation and integrity of information contained in the memory of the mobile device and its components. That is, they're divided into software measures, hardware ones, and for powered-on devices and for powered-off ones. It spells out which methods are used to do this. Recommendations are given on diagnosing faults in a mobile device and restoring it to working order. Actually, in our very first methodology there was also a tiny little table devoted to faults like these. Now it's a serious appendix in which all faults are broken down into five types. The device doesn't power on; the image on the display is partially or completely absent.
The device doesn't respond, or responds incorrectly, to the control buttons. No response from the device to touchscreen input. And no communication with the device through the connection port. Including when the device fails to initialize in the operating system of the bench equipment. For each of these items we've laid out the possible causes of these faults. The table's second column, and the third — what to do in each case when such a cause has been identified. Sometimes, of course, it turns out that's it, the device — nothing more can be done, a brick — so write it's a brick. But that's only one or two items. So this part has been added and nicely, interestingly expanded.
Next. What should we do as far as establishing the passcode on the device, the PIN or PUK code, or the SIM card password? In the two previous talks we were cracking the password. This way, that way, that other way. But what if that doesn't work? What's needed in that case? What other options do we have to try? Well, actually, we have the motion. And the investigator can question the person from whom this phone was seized. Various methods — actually, the operational way of obtaining password information hasn't gone anywhere, and operational combinations do exist. So, among other things, this methodology, besides cracking the password by various means, establishes the types of motions: for provision of the password value for access to the mobile device's memory, for provision of the PIN or PUK code values for access to the SIM card memory. And the coolest one — for the device's user to be present at the forensic examination in order to pass biometric identification, because a number of actions with various phones, are sometimes impossible without that.
So these are provided for, and templates are given for how to do it. If they don't provide it, they don't — the expert often can't wriggle out of it, but we've provided for these things. Next: the methods and stages of extracting information from mobile device memory — these have also been worked out. What's included here? As for methods, it's low-level extraction, meaning full file system access — that's one. And second, extraction of publicly accessible data. So these methods exist, and sometimes the first works, sometimes, unfortunately, only the second does. And the stages. These aren't stages the expert must necessarily go through — one, two, three. It's what he may encounter with these methods of information extraction.
Well, like, powering the device on and off, photographing, video-recording the screen — that was the very first to appear back then: what do we do when we have to extract information from a mobile device? We had to, generally speaking, take a camera and, swiping the information with a finger, photograph every single screen. So that hasn't gone anywhere; some things still need to be photographed too. Connecting the device via a connecting cable or wireless interfaces to the expert's workstation. It includes establishing and initializing the connection between the mobile device and the expert's hardware and software workstation.
Putting the device into a service mode using test pads or the control buttons. Installing an agent program, downgrading the operating system and software version, bringing the device to a state that allows information to be extracted from it, connecting the SIM card to the expert's workstation using a SIM card reader device. So, let's say, the expert includes and goes through a number of these stages in an examination. The next thing we got again — well, the glossary of key terms and definitions has been significantly expanded compared to the methodologies that existed before. Every methodology has a glossary, but we've expanded it a little. And so on — there's no point listing everything, but quite a lot specifically about how to examine the device has been added; all these procedures are there.
What else is very interesting? Well, let's say, there's a problem we have and have always had. And the standard methodology for examining computer information says that, in principle, the resulting information should be copied to write-once storage media — everyone will have fewer problems. The investigator — how to read it later — and the expert has fewer problems too, but we're coming to the point where data volumes are growing and write-once media are no longer sufficient. So this methodology now includes regulation of exactly how the sought information is written to rewritable storage media. In what form? If the results of the examination are written to a solid-state drive or an external hard disk drive, then the explanatory label and the expert's signature are put on adhesive paper tape attached to a part of the drive that has no identifying features: the serial number, some other number — none of that.
A corresponding entry is made in the research section of the expert report, stating the identifying features of the storage medium — well, like the make, model, serial number and so on — the size of the written information in bytes, the cryptographic hash value of the information written to the storage medium. Then at least it'll be clear how to work with these things and how to write to rewritable storage media from now on. What else is new? And now the most interesting part, because this was a very problematic question, and we even held a meeting with representatives of the EKC MVD of Russia to work out a single common approach. What do we do if, in the objects under examination, we find credentials for authentication on cloud services, cloud storage, passwords for email and so on? I mean, we were literally tearing each other apart: what, we found it, we have the question, and the investigator needs it.
So they got into the email account, downloaded all they needed from there too. Got into the cloud storage and downloaded all they needed from there too. Well, there were, like, several sensible objections that, on the one hand, this is probably already exceeding our authority, because now, now, now — now, there's such a thing; I was against it too, but unfortunately, many think so. And on the other hand, it's also an additional large volume of information that the expert has to process, given that the backlogs there are huge, and working through this too... My mailbox has over 10,000 emails in it, and you wouldn't much enjoy working with it. Go ahead, pull down all the rest — little of it would actually be needed. In any case, what did we decide? First: establishing the presence of data for authentication on cloud services and cloud storage — the so-called tokens — is a stage of the information examination. We've now officially put that into the methodology, and that's important.
And then the following was decided. If the mobile device's memory contains files with authentication data for cloud services — tokens — and actually, I suppose this would also cover, and I'm thinking we might still manage to add, information for accessing cryptocurrency and other things. Although they sometimes ask for that info, and there the keys got listed fine — on internet resources, in social networks, in email — immediately, without waiting for the examination to end, once you've found it, write a notice about it to the investigator, attach that information — well, obviously — and then attach it to the expert report and note the possibility of extracting data from cloud services later on, in the course of investigative actions. So you report at once; the investigator, if he needs it, comes and inspects too. The main thing is he picks out what he needs in the given case.
And a standard template of such a notice on the presence in the examined objects of authentication data for cloud services and storage is provided in our methodological recommendations. Well, let's say, one more problematic issue that actually hasn't been resolved anywhere at all.
Much as we'd like, so far nothing seems to be said about it in Federal Law 73-FZ — yes, nothing at all. So what do we do when the investigator simply writes with mistakes? I'm not even talking about him framing the questions methodologically wrong. We don't go into that; we simply rephrase, saying the question is put incorrectly, that it's a legal question, and so on. That's not what this is about. But lots of our experts do the following. They put the questions in quotes and copy them verbatim, with spelling mistakes, with syntax errors. And why should I have to correct them? But I understand: on one hand, the expert himself may not be very literate either and may add his own mistakes in place of those.
But we resolved it this way. If the questions put to the expert contain semantic, stylistic, lexical, syntactic or spelling errors, then in the introductory part of the expert report the question wording may be edited, indicating the reasons for making the corresponding changes. So they may be — if you don't want to, don't consider yourself literate enough, you'll write it as is — well, that's one more way to show the judge that the investigator isn't very literate, and what are we going to do about that. So in fact that's the option; we decided to put all of this, these provisions, into our standard methodology.
And so, at a meeting of the Scientific and Technical Council of the Investigative Committee of the Russian Federation, this methodology was reviewed and basically recommended for use in the forensic expert units of, first of all, the Investigative Committee, then we'll publish and distribute it — well, distribute it and, in any case, bring it to other forensic expert units. It's their right to use this methodology, to cite it in their expert reports, or not to. But we've already given it a certain direction for its use. Right now I have no printed version of this methodology, because it was sent to the proofreader and then on to the typesetter, and it will be printed. But it's not a quick process. Given that at the end of the first half of the year we held the Scientific and Technical Council, at best — at the very best, it comes out in printed form by the end of the year. But we're already announcing it and drawing attention specifically to which problems and tasks we set ourselves and how we solved them.
Thank you. —
[applause]
— So, colleagues, your questions. Mikhail Mikhailovich, on my way. —
— Olga Vladislavovna, thank you. Well, I can't keep quiet. So, first. A harmless one first — just to clarify video recording of the expert's actions. That's probably connected with particular cases. I remember some of them, but please explain to the audience. —
— No, actually, sometimes there's a need to record what you did, where you went and how it happened. So why not? And not only substitution — there were complaints when they say you swapped something, did it yourselves — that too, probably. No, well, it's for when the need arises — you see, these are all advisory, recommendations in general. We'd like you to have all of this, and when you start working with it, if such a task comes up, that you have the equipment and can record all of it. But as a rule, video recording is still needed more during inspections than during, say, an examination. I'll tell you a real case: a phone comes to me, it says a Samsung phone, I open it — a totally different model. And the courts say, was it video-recorded? So the fraudster brought to court a phone on which he'd committed the fraudulent actions, but swapped the phone itself.
Well, the judge just blindly recorded it — they don't know which phone it is. And that was a real case. No, well, look, I also brought up feature phones, and a number of feature phones — old feature phones — could only be examined by setting up a video camera, a still camera. Now here's another thing: as I understand it, you didn't include smart TVs and smart set-top boxes, since you'll have some smart-home methodology, and they'll be assigned there, or how? No, well, at one point one of our respected developers made us equipment for examining them, but practically no examinations were done with it by connecting to anything. —
— Well, these days I easily go online from my TV, and video and everything else. Right, but those involve other examination methods and methodologies. So that would be some separate methodology? If such a need arises — that we're constantly having phones, I mean TVs, brought in for examination — then we'll have to develop and write such a methodology. No, in fact they hardly ever bring them. Which is why that equipment sits with me — honestly, it's sitting on the desk. But unfortunately, so far I don't see any particular need for such examinations. —
— And one more — okay, I'll end on this, I understand. So, look, regarding data that is personal: the personal data laws, among other things, allow third parties, responsibility for data lies with others, and the expert, under Federal Law 73-FZ, also bears responsibility for disclosure. So I don't understand what the problem is with presenting this data, especially if it's required. Well, as I understand it, you — I mean those tokens, what you were saying. No, there's no problem at all presenting them; it's just that the order of how to do it, and the procedure, we've spelled out. When I raised this question for myself and tried to talk it over with everyone, I even ran a little survey on Tsifropol; some there said, yes, yes, let's examine it right away as part of the examination; some of my experts, whom I still keep in touch with at the MVD, say, well, we too sometimes turn a blind eye and hand all this over in the examination.
But it was the EKC MVD department heads who said, for God's sake, we don't need this, because it'll be too complicated later for every examination. It'll complicate all this work, add to it, and our backlogs will grow again. We just don't need that in such a format. Thank you. —
— So, colleagues, more questions. There, I see a hand. —
— Hello, wonderful talk. I'm always impressed at every conference when you come and tell us about the methodology. I have a question regarding cloud storage. For example, a real-life made-up case: examining a computer running Windows, and it's syncing with OneDrive. The sync covers official documents, including documents and information constituting state secrets. We see that the sync has happened and the information has leaked to the servers of a foreign state. How should the forensic examination be conducted and written up in this case? —
— If only I'd caught all of that. —
— On the question of examining classified information, you know, well... No, the situation is that the computer is syncing with OneDrive. That's standard practice on Windows systems, if it hasn't been turned off. And the sync happened with documents — official ones, for example "for official use only" (FOUO), of various kinds. How do you write up the forensic examination in this situation, if we know that these documents have leaked to the servers of a foreign state, and OneDrive falls under laws like FISA or the CLOUD Act, which... The problem is that such documents shouldn't be on such computers at all. That is, a person's internet — I mean, a computer that goes online must not contain any FOUO or classified information — none whatsoever.
If it leaked, of course, you must report that such information is there, because it's simply a violation in itself that such information exists on such a computer. Second, if you see that it leaked, you absolutely must report that. Here there's not even any disagreement in understanding that, yes, these documents are also sitting, among other places, in cloud storage. But as a rule, still, let's say — inspecting cloud storage clearly has to be part of an investigative inspection. So, as part of an inspection. That inspection procedure exists. With video recording, naturally.
Yes, video recording, absolutely. But actually, our biggest problem isn't that; it's whether our experts are cleared to work with information involving some state secret. —
— Well, let's say they do. However many times we raised this issue back in the day, we were told: well, you must have a specialized hardware-software system that is certified, tested and so on. And we say, wait, wait, that's all fine for working with such documents, everything we have is certified and so on. But when they're set up so that no other documents can get into them — files that aren't meant to be processed on that computer — that's a very serious problem, and we can say that there's a file with such-and-such attributes, such-and-such markings and so on, and nothing more. But unfortunately, we have no other option. "Possibly containing." Always "possibly", right? —
— "Possibly containing." Well, that's how we write, right? No, we don't even say "possibly containing". We say there's a corresponding classification marking. Actually, I know perfectly well that if a person needs to keep everything on a computer — I'm not talking about anything classified, even just an FOUO document — he simply goes and wipes out the relevant numbers, markings and all the rest, and keeps everything else. And you'll never determine whether it's classified or not. Thank you. —
— Right, colleagues, one more question. Nikita, over there, I saw a hand go up. —
— Yes, San Sanych.
Tell me, please, about the Scientific and Technical Council. Is it working on describing and creating a standard mobile forensic laboratory that would bring together all the main tools, like those from MKO, Elcomsoft, ELETEK? And are there any thoughts, let's say, about methodologies? Thank you. —
— So, the Investigative Committee's Scientific and Technical Council — like, in fact, any such body; we have them at the Russian MVD too, and I think at the FSB — everywhere there are certain bodies that review and approve; they don't develop anything. If I, as head of a certain research department, together with the experts who work with us, with the Investigative Committee Forensic Expert Center — we developed this method, we have the right, and we refer it to the scientific and technical councils, which then recommend it. So it's advisory in nature. What you're talking about gets developed, let's say, not within some scientific paradigm; as a rule, it all goes as an annex to a state contract.
And it would take you a very long time — as I say, today we have one vision of this equipment, six months later another; say two companies started cooperating, and now we have a different little label, a different name for the software everyone uses, and then they went and fell out, and now we've got two different labels again, and we need both of them. In general, actually, we always have a problem between what we want — and what we want is for it to be one, two, three, done. Software products partly have the same functions, but this one has one bell and whistle, that one another, and I need both. And it's always very hard when you're allocated a specific sum, no more, for procurement — to fit into it with your bells and whistles and justify that today I want Mobile Criminalist, and tomorrow, somehow — what if Cellebrite comes back, then I'll want Cellebrite.
It will all depend on the market and what's available. What you mentioned, the development of systems — yes, that's basically a separate research effort that's commissioned, and within the MVD it's easier to do: they have a special unit, "Special Equipment and Communications", that handles contracts. I acted as the functional customer. I'd say, I want to develop such-and-such a hardware-software system now. Functionally it must include these things. And our contractor, who enters the tender, develops both the description and the designation, and presents a prototype. We test it. And we say, all good, great, it's adopted into service with the Russian MVD. A year later everything's different — well, not everything, some of it; they come and say, let's refine it now, change this here, change that there, now your system will have this configuration.
Well, let's change it — tests again, all the rest. We say, yes, this suits us, now we'll be procuring equipment with these specifications. That's not science — it goes specifically through development (R&D) work. —
— Olga Vladislavovna, thank you very much. Let's send her off with a round of applause. That was really great.
5. Alexey Moskvichev (MKO Systems) — “Applying forensic methods to investigate thefts committed with NFC technology”
Scheduled 13:05–13:35.
Moderator's introduction
So, friends, by and large we're moving on. And before we move on to the next topic, let's chat a little. Surely, it seems to me, everyone has some favorite sound. Well, I like the "ding" when, for example, my salary comes in. But then again, that "ding" may not always be a good thing, because the "ding" when the salary comes in — that's good, that's happiness, joy, dopamine. But sometimes you also get that "ding", and money is debited, and unfortunately it's not being debited to you. It's exactly such cases my colleague Alexey Moskvichev will talk about, and how they're investigated. Please welcome him with applause. Alexey, the microphones, everything is yours.
Talk and Q&A
Can you hear me? Yes. Dmitry, thank you very much. Let me start by introducing myself: at the company, I'm in charge of training. Apart from that, my colleagues and I regularly travel to the regions, we run a series of seminars. And at one of those seminars, some information was shared with us by our wonderful user, who I hope is watching us, he couldn't make it here, unfortunately. That information became the basis of my talk today.
Well, before I get to the substance of the announced topic, let's briefly talk about what NFC actually is and go over its key features. So, NFC is a short-range wireless communication technology that uses radio-frequency identification to recognize objects and devices and to read information. Basically, it has become a firm part of our everyday life and is widely used in mobile wallets, access control systems, for example, access to office premises, and in most self-service systems. And here in front of you are some of its key features. First, it's short-range communication, that is, NFC's range is 3 to 4 centimeters, which, on one hand, ensures its security, and the connection's precision. Next is fast pairing, that is, it's enough to simply hold two devices up to each other for a connection to be established.
The next feature is low power consumption, that is, NFC consumes very little energy, which makes it convenient for small devices that in some cases need a long standby period. Basically, it's supported by most general-purpose applications, as I've already said: mobile banking, personal identification, transferring data, photos, contacts, paying for public transport fares, and much, much more.
But who would have thought that, on the one hand, this seemingly convenient and supposedly secure technology would become a serious weapon in the hands of criminals, one capable of draining the personal accounts of people all over the world. And here's a bit of statistics, global for now.
In 2025, the number of crimes committed using NFC technology grew 35-fold compared to the second half of 2024. Just think about that figure. Of course, this was helped along by the emergence of all sorts of malware and new relay schemes. And in particular, listed here in front of you are several such applications, using one as an example, we'll go through a practical case from a real criminal investigation. The first is the NGate, or NFCGate, application, which transmits NFC data from payment cards via compromised smartphones, for subsequently carrying out fraudulent cash withdrawals at ATMs. The next application, GhostTap — behind it are, by the way, Chinese developers — steals card data, and loads it into digital wallets for making contactless payments.
Basically, phishing is used first, and then the payment card details of the victim are linked to the suspect's device. And one such device can have from 4 to 6 sets of card details linked to it. Later, on the secondary market, such devices are sold via various Telegram channels, even for hundreds of dollars. A similar application, SuperCard X, also, by the way, has Chinese people behind it. I've said "Chinese" twice now, draw your own conclusions. Likewise, on the one hand, it presents itself as a supposedly safe application, but on the other, it actually covertly collects bank card data and transmits it for conducting quick illegal transactions.
And a huge number of Telegram channels are springing up like mushrooms, publishing detailed instructions, training videos that even untrained users can understand. In particular, while preparing for this talk, we looked through some of those Telegram channels. Some of them, by the way, have thousands of members. And what's also notable is that from the context of the chats it became clear that some of the groups that use these fraud schemes are targeting, for example, the United States, the UK, Australia, Canada, others target Malaysia, Japan, Taiwan, and still others, African countries. That is, as you can see, there's a certain division along regional lines. And let's keep Telegram in mind here, because we'll be coming back to it.
What about Russia? What about us? We've got gas in our flat, as the rhyme goes. In Russia, the first reports of crimes committed using NFC technology appeared in August 2024. Back then, the amount stolen was around 40 million rubles. And, well, analysts predicted such crimes would grow by around 30-35% a month.
And basically, that's what happened. In 2025, around 400 such crimes had already been recorded, committed mostly using the application NFCGate, or NGate, and derivatives of that application. And the average amount stolen was around 100 thousand rubles. And, basically, two criminal schemes were used. Let's briefly go over them too. The first is the classic scheme. The victim's phone receives a call. The fraudster convinces them to install some kind of "specially secured" application on the victim's device, after which the victim is asked, following instructions over the phone, to hold their bank card up to the NFC sensor of their smartphone, to enter the PIN; they insist it's safe to enter the details, the card itself stays with you, the PIN isn't dangerous, but in fact this very application collects the NFC data of the victim's bank card and sends it over to the suspect, who at that moment may be standing at a payment terminal, at an ATM, and holds up their own identical device with the app installed. So there are already two devices here.
NFCGate on the victim's side and the same application on the suspect's side. And basically the money gets cashed out. The essence of the reverse scheme is that some kind of malware is installed on the victim's phone, which relays the signal from the suspect's card to the victim's device. So in fact the sequence of actions is different here. The victim is guided over the phone to an ATM and asked to deposit money into a supposedly "safe account", but it ends up on the suspect's card. So those are usually the two schemes. Let's go through how NFCGate, or NGate, works.
In general, it's a legitimate application for capturing, monitoring and analyzing NFC traffic by intercepting and replaying it. It was actually developed as a student project at one of Germany's technical universities, in Darmstadt.
The source code is on GitHub; by the way, you can have a look at it, if you're interested. See, there's a QR code here on the right, you can go and study it in general, if you're interested. And in fact this application has several operating modes, each of which handles this NFC traffic in its own way and basically works with it differently. Here's the first mode, Clone Mode. It reads a tag and instantly replays its signal, but here, to replay this NFC signal, a rooted device is required, that is, a device with superuser rights, in order to replay this NFC signal. For example, when could this be useful.
Say I need an access card, either to an office or to a warehouse. Somewhere at the reception desk or, say, in the parking lot next to the office, I try to read, using a device like this with NFCGate installed, some employee's card. Then I make an exact copy, and basically I can walk into the premises unhindered. That's the essence of this mode. The next mode, Relay Mode, lets you transmit this NFC signal over the network. That is, in this scheme, several devices with NFCGate installed are used. That is, one device acts as the reader, meaning it reads the tag. And the second device acts as that very tag, replaying the signal.
Here's an example. Somewhere on public transport. I have a reader like this with NFCGate. I try to read the victim's cards. At that moment the card may be in a pocket, in a handbag, doesn't matter. Meanwhile my colleague, my accomplice, is anywhere in the world at all, and picks up this signal using his device, which acts as that very tag. And then he can make, for example, contactless payments. But again, there are nuances here. For example, in Russia there are limits on contactless payments. Most POS terminals support them, and on most POS terminals that limit is 3,000 rubles, on some it's 1,000 rubles. So you probably won't get rich, but probably enough for gum and a Coca-Cola. Here you'd probably have to go for volume.
The next mode, Capture Mode, is mostly, I suppose, geared towards penetration testing. A pentest, right? That is, passively collecting this traffic to see whether or not the NFC signal can be replayed. That is, looking for vulnerabilities and possible ways to break in. This mode is geared mostly towards that. And finally, the last mode, Replay Mode. We've already said that NFCGate is an application for intercepting and replaying NFC signals. So in this mode, you can emulate a bank card over and over again. That is, having recorded once, recorded this NFC traffic of one of the devices that were communicating, you can emulate it repeatedly, that is, emulate the bank card an unlimited number of times.
That's the essence of this mode. So, since NFCGate first started being used in attacks on bank customers, this app has seen a number of modifications. And in fact the fraudsters themselves have learned to disguise these applications as popular government or banking apps. And in fact this new fraud scheme caught the security staff of credit institutions off guard, and what was found was more than 100 unique applications, or derivatives based on NFCGate, which, basically, handled NFC data in different ways. So let's go through how attacks using NFCGate work, as an algorithm, in its pure form. What, in fact, does it start with. The first stage is most often based on plain social engineering, that is, the victim, under some pretext, renewing a mobile contract, hacked Gosuslugi, bank card protection, renewing a health insurance policy, it varies, all sorts of variations are possible here, is asked to install this certain application on their device.
And this application looks similar to a legitimate app, either of a government agency or a bank, but in fact it's precisely NFCGate, which most often runs in Relay Mode. As for remote installation of this application, so they've supposedly convinced them over the phone, and contact is established. For remote installation, so-called RAT applications are most often used, or remote access trojans, which most often arrive via a messenger as APK files.
In the NFCGate being installed, as I said, Relay Mode is already running. The server settings are specified, where the NFC data is sent when the NFC tag is scanned by the victim's own smartphone. And at that moment, the suspect's device is likewise running an identical application in Relay Mode. And he's basically waiting for the victim to launch it, so he can move on to the next stage. And so, basically, the device, here are the screenshots in front of you, on the suspect's device NFCGate is running in the so-called Tag role. Hard to see, but here's the Tag role, and on the victim's side, the Reader role.
And how, after all, is he to know he can already move on to the next stage, and that the victim has actually launched the app, and not just stringing him along over the phone, so to speak, saying "yes, yes, I've launched it, all good, we can move on." Essentially, he sees that this session is open — you see, here's the little green one, it's hard to see there, the label signals to him "Network Connect to Partner", so the connection is established correctly, we can move to the next stage. And the next stage of the attack is reading the card data, intercepting this traffic during authentication at the ATM over the NFC protocol.
And essentially, at this stage the attackers are, as I've already said, next to the ATM and hold their own device with NFCGate up to the terminal, and all that's left is to enter the bank card PIN. And here, after all, how do they also get the PIN from the victim's side? Well, essentially, it can be, the same old social engineering can be used, or virtual keyloggers, which record the victim pressing the virtual keys, right, and using those same RAT applications they send this information to the suspect's side. In some NFCGate modifications there pops up an additional window that requires entering the PIN. So here it's already the victim himself, who — well, most often when using those apps that mimic banking apps.
So here it is effectively the victim himself who hands the data over to the other side. And after a successful scan, essentially, the suspect gains access to the personal account and can withdraw the funds.
The next stage, well, the final one rather, not the next, is the reuse of this intercepted traffic using Replay Mode, as we've already said. That is, re-emulating the operation of the bank card. As I've already said, the new wave of attacks using NFCGate most often happens by distributing APK files that emulate the operation of legitimate apps, either a government agency or a banking app. And there's a common pattern observed when dealing with such applications. That is, most often we see the UI being changed by creating a stylistically similar graphical interface. The fraudsters try to hide push notifications so as to effectively prevent detection and a timely response by the victim to the emerging threat.
Next, we see changing app package names, changing the format of the collected data. Well, essentially, as I've also already said, in some modifications an additional prompt pops up to enter the PIN for the victim's bank account when the app is launched. Well, and essentially, let's go through the theory in practice. Okay, a real case. In this modification the mode was preset, the NFC operating mode, to Relay Mode. After that the data was sent to the attacker's server. Here, essentially, is the scheme in brief. A call came in via Telegram. An unknown person introduced himself as a law enforcement officer.
And under the pretext of keeping the money safe asked the victim to install a certain app. Then he gave instructions on how to launch it and what to do with the bank card. Then he recommended deleting this app. And as a result, funds were withdrawn from the victim's account in the amount of 230,000 rubles. And what came in for forensic examination was a Redmi Note 11S mobile phone. Well, essentially, let's go through the steps the forensic expert took. To begin with, using our software, Mobile Criminalist Expert Plus, an advanced file system extraction of the device was performed; specifically, the MTK Android method worked successfully.
Since, as we've already said, these apps are detected by most antivirus applications, right, we performed a scan of the file system structure of the extraction, and, essentially, a file called vtb1 was found, which was detected as a trojan. Essentially, looking at the directory, you can see it was created, modified on 22.01.25 at 10:20. And from the directory you can see that this file was found in a directory associated with Telegram. We can assume that it, most likely, arrived on the device in the course of a chat via the messenger. Moving on.
Essentially, when analyzing the contents of a database table located in the directory associated with Telegram, you can see that this vtb1 file was received from a Telegram user with ID 7 and so on. We've hidden part of the identifiers because I'm not sure how unique they are. Well, still, when it's a real case, to rule out any complications, I'll just comment on it. In the next table, note that we — the expert — found the identifier of the downloaded file, with the value 52, then 93.
Then, by this identifier, 52, 93, when analyzing the contents of another database, cache4.db, also registered in the directory associated with Telegram, the date of receipt of the message with this malicious file was established. Note: 22.01.25 10:20:44.
And if we go back a step, it matches the modification date of this file in the Telegram directory. Essentially, once it became clear where and how this little beast came from, all that was left was to see what it actually is. Well, the next stage. After decompiling the vtb1 file, in the manifest XML the application identifier "Darmstadt" was found. Note, it isn't highlighted for us.
Let's try.
Here it is, the identifier of this app, if you can see the pointer. The name for this identifier is also displayed in the AppName parameter, VTB Protection, in the XML file as well. And we also identified — we've masked it too — here we identified the IP address of the server that the intercepted NFC data was actually being sent to. What else is notable here? Further examination of the trace picture this app left in the smartphone's operating system showed it was uninstalled from the device twice, on 23.01.25 at 11:16 and 11:41.
And the install history of this app is also fully consistent with the timeline, starting from the date the malicious file was received. Next, in a directory related to the device's file system, there's a record of the NFC system service starting, which also indirectly confirms that this technology was in use on the device. And finally, in a directory related to mobile banking, we found an image file with a record of a withdrawal of 230,000 rubles, which, I suppose, confirms that the theft of the funds was completed.
And in conclusion — prevention matters too, I think — we suggest a few simple rules that help protect against carding, as it's called. First of all, you should install apps only from official stores; don't share your bank card details with strangers and don't enter them on any suspicious websites or in apps. If you get a link to install or update a banking app, the first thing to do is probably to call the bank's hotline and check whether that offer is genuine at all. And if the bank card has been compromised after all, try to block it as quickly as possible, likewise via the hotline, or through the official banking app.
To finish, I'd like to quote the greatest swindler of all time. But not all of you know this film — "The Twelve Chairs". The financial abyss is the deepest of all abysses — you can fall into it all your life. So may your service be the anchor that keeps ordinary citizens from falling in because of vulnerabilities like NFCGate and its derivatives. Thank you.
[applause]
Dima, there's a colleague up front here. Alexey, you wrapped up awfully quickly there. Thanks for the talk. My name's Alexey too, by the way. Here's my question. Does the NFC app leave digital traces when it operates in Clone Mode and Relay Mode? I mean, you said — when they walk up to someone. You've just described the traces of what the victim installed on their own device through Telegram. But if I understood you correctly, in NFC Clone Mode, the attacker just walks up to the victim, right? Not as such, no. It's more about the trace side of operating system artifacts, of the NFC technology itself starting, but there are very few of them.
So they don't leave traces on the device? Effectively, no. And in Relay Mode, when there's an intermediary? No, Relay Mode is exactly this case. Yes, here there'll be one — the trace picture.
— Relay Mode, there will be. You also mentioned Replay? No, Relay. Relay. In Relay there will be. In Capture Mode there won't be — there'll be very few. But in Relay there definitely will be. —
— Those modes — do cases with those modes come up at all? Any practice out there, heard of any? Honestly, it was a real stroke of luck that we got a case like this at all. I think — someone may correct me — there aren't that many. And, in fact, probably not that many devices end up submitted for examination. Maybe someone here has a different experience? No, we get plenty of APK files. Where I'm from, we get them all the time — at least once a month for sure. —
— No, but actually... More often, actually, we get victims coming to us who installed an APK file, and then afterwards the attacker told them to delete it, and then we try to go down the same route through the EKC. Well, I think if you have the victim's device, you'll definitely find the trace picture. I was just curious about this particular case. In Relay you'll definitely find it. Clone Mode. Thank you. —
— I see a hand, on my way. —
— Hello. Two quick questions. You said a reader is used — that it could be used on public transport, for instance. And you also said NFC works at a range of 3 to 4 centimeters. Is that range enough, on public transport, say, to read the data, with that many people around — that's one point. Or maybe there are some antennas, amplifiers — so that's question one. —
— Well, we probably don't have that much expertise here in that respect. We didn't do that examination ourselves — this is information that one of our users shared with us. But judging by the information we have, in principle, it is realistic to read the data. So that data, even without a PIN code, is enough to carry out some sort of... To read the tag, yes. —
— And to carry out some transactions. Effectively, yes. —
— Thank you. That's one. And the second question. You also touched on SuperCard X, that it's positioned as legitimate or actually is legitimate — did I get that right — as a contactless payment app, that is. —
— That's right. So how do you tell, in that case? You said it sends... A range of modifications, only here based on that app, just like the ones based on NFCGate — its derivatives. Ah, so it's not the original app? Not in its pure form, of course. NFCGate too... Built on it. Yes, of course. Thank you. —
— Right, Nikita, could you go over, please.
Good afternoon, thanks for the talk. I have a question. You mentioned that besides Relay Mode, Replay Mode is also used in attacks. Can you give any examples of this type of attack? —
— Once more, louder please. Besides Relay Mode, Replay Mode is used for attacks. Are there any examples of actual real-world attacks using that mode? —
— Well, we don't have that information. Yes, we can't give an example in that sense. But essentially, like I said, it's the same mode of retransmitting that signal multiple times. That is, having recorded the data once, you can use it over and over.
Sorry, in that mode, as I understood it, there's also the security code, generated fresh for each transaction. How can that be reused afterwards? No, not the security code — the bank card PIN. No, no. I mean the security code specifically: when we pay via NFC, a new security code is generated every time. Ah, the one-time one? Yes, the one-time code. So how did you copy it, when it's different each time? How does that work? Well, that only works via social engineering, when the victim is on the phone. And then it works.
Right, I see a hand there.
[music]
What a shame there are no users in the room, right? The roast hasn't even started yet, and the battle's already about to kick off. —
— Thanks a lot for the interesting talk. A question about the antivirus. It detected it on the workstation where the examination was done. If it had been running on the phone itself at the time, would it have detected it on the phone, or would it have missed it because of a sandbox or something? —
— It would. It would on the device too. —
— Okay, thanks. —
— Colleagues, more questions. Up front, Nikita. —
— Where? Artem, as far as I remember. —
— We know our regulars by sight. Good afternoon. SQL question. As I understand it, you parsed all those values in the SQL databases by hand.
Do you have a reference guide anywhere, with a list of all the files and what's stored in them? Like, this file is responsible for this, this file stores this kind of data, that file stores that kind of data. Because we also have to search for a lot of things. Manually. —
— Some of that information is available, but for this case it was manual only. The forensic expert did the work, I gather. —
— So that's the question. Could you release a reference guide like that, with information about which data is stored in which file in the apps?
We'll try; maybe we'll add it to the list of requests. The most honest answer. We'll try, but we're not promising. Nikita, did you see the hand there? Definitely, yeah. No, we're preparing a lot of things this year in general. The speaker's feeding back, I'll step aside. We're preparing a lot of materials, but we'll try — no promises. —
— Good afternoon. Hello. Here's my question: is it possible to copy an NFC tag, say, on a device, on one Android device, let's say, with installed… Could you raise the mic a bit, please, it's hard to hear. …payment apps, like Mir Pay or a payment sticker, and the attacker clones it with a second device?
In theory it's possible, but in practice we haven't tried it. —
— Yuri Mikhailovich, there's still the roast, the time for battles, let's leave it for later. Go ahead. For the roast? —
— No, why? One more question is fine, but if a fight breaks out now, it'll be on you — and on Nikita. —
— I just wanted to clarify. You showed a slide where an APK file was scanned with Kaspersky, specifically. Was that the Kaspersky built into the MKO software, or a separate standalone one? It was separate, yes, but we can do this inside the product, in the "Malicious Objects" section. So now this is our integration with Kaspersky, so you can… Unfortunately, I've seen cases where it doesn't always… —
— Anyway, those databases are up to date in the customer portal, so update them. That's the only advice I can give here. Thank you. I don't know why Alexey looked at me. The database is up to date, update it, all good, we check it. OK, colleagues, the rest of the questions go to the roast. Let's give Alexey a round of applause. Alexey, well done. Thank you very much.
6. Igor Zaitsev (Account-Best) — “UAVs on the line of contact and in rear areas. Methods of use and countermeasures” — not in the recording
Scheduled 13:40–14:15. Closed-door talk, not shown in the stream.
Moderator's announcement
And before we move on to the next talk, a short announcement for our online viewers. Unfortunately, we can't show the next talk in the stream, and we're back only after the break, which will be at about, well, it'll be 2:30, around 3 o'clock. So, and we're moving on to the next topic. The next topic, actually...
[Igor Zaitsev's closed-door talk and a break (13:40–15:00 in the program) are cut from the recording; timecodes run without a gap.]
7. Sergey Eremin (LAN PROJECT) — “Using VR-EXPERT to examine DVRs. A comparison with foreign counterparts”
Scheduled 15:00–15:30.
Moderator's introduction
I invite everyone to come back into the event hall. We're starting our third part now. You can bring along whatever you haven't finished eating, that's no problem at all. We'll finish it all right here.
I'll say right away that our curtain is about to close.
Testing, one-two — and like this? OK, anyway, let's just carry on. We continue with our talks. In today's world, on the road there are almost always small witnesses with us that we probably don't even notice. These witnesses are stuck under our windshield, always amassing information, collecting it. But, again, the question is how to process and analyze that information afterwards. Sergey Eremin from LAN PROJECT will tell us about that today. Let's give him a round of applause.
So, Sergey, the floor is yours. The mic, the clicker is right here too. Go ahead. —
Talk and Q&A
Hello, dear colleagues.
My name is Sergey Eremin. I represent LAN PROJECT. At the start of my talk, I'd like to note that our company turned 25 this year, and throughout those years we've been developing hardware-software systems and software solutions for acquiring and analyzing digital information.
We supply all law enforcement and security agencies, and we hope that the products supplied by our company satisfy you in terms of quality and the completeness of the information you obtain and can use in solving and investigating crimes, as well as in detecting them. And today our systems and our software are aimed at various kinds of software — or rather, digital information objects, whether that's mobile phones, personal computers, or cloud data of some kind. We cooperate with many vendors located in the Russian Federation, and we develop various solutions together with them.
We also have extensive experience cooperating with foreign vendors. Yes, under current conditions, supplies of foreign software are difficult, but at least we see what's going on in the Western market and know their capabilities. And, where possible, we also try to implement some of those capabilities on our market. I'd also like to note right away that, to mark our company's 25th anniversary, we'll be holding an additional prize drawing from our company. To take part, please come to our booth, register separately, and after the event ends there'll be a drawing, and the lucky winners will get prizes from our company.
In today's talk I won't cover the full range of solutions we have, that we've created; instead I'll focus on one niche that we've been working on very intensively for two years now — a niche in which, on the Russian market, no equivalent Russian software existed before. That's video recorders — a large class of devices, video recording devices, both stationary and in-car. During the talk I'll go into this in more detail and tell you what we've managed to achieve and what we're planning in the near future.
That's not my presentation.
So, today I'll talk about the functionality we've already implemented, what we're planning, and we'll briefly try to compare our software with foreign counterparts. First, for those who are seeing my talk for the first time and aren't familiar with our product, I want to say that our software is called VR-Expert, and this year it was registered in the Russian software registry, which makes it fairly easy to supply to government agencies.
We essentially have two main areas we work in, which are already fairly fully implemented, and we're also working on dashcams — I'll tell you what's been done there and where the problem lies, too. Well, stationary surveillance systems are a fairly widespread type of object — in Moscow, perhaps less so, because the Safe City system there is implemented very well, practically all video systems are tied into one network, and getting information from them is fairly easy. But in the regions, Safe City isn't so widely developed everywhere, and you quite often come across standalone DVRs in various offices, institutions, and perhaps homes.
And these objects are an additional source of operational intelligence and evidentiary information. But the main problem with them is that we can't always get access to it. What's the reason for that? Practically all stationary DVRs have their own proprietary data storage system, which ordinary operating systems don't recognize. And most often — very often — we've run into situations where a disk comes in for examination, the disk gets connected by whoever received it — either a forensic expert or some officer the disk was sent to — who figures they'll just take a quick look at it, and plugs it into their workstation.
A cheerful window pops up, saying, "I've found a disk here, let me initialize it for you." The person clicks "well, yeah, initialize it" — what happens after that? The information becomes inaccessible. Whereas before, back with Windows XP and 7, only a fairly small part of the boot sector got overwritten, now Windows 10 and 11 stuff roughly the first forty megabytes with their own information. And in many cases, on most DVRs, it's exactly this area that holds all the metadata and the allocation table showing where the video streams are located. After all, essentially, most recorders on the market have no file system as such at all. That's why Windows doesn't understand it either; the recorder just knows which disk sector its video stream is stored in, and which disk sectors hold the markers indicating where each video stream lies. That's exactly what's needed to decode this data.
And what's more, every manufacturer builds this structure their own way. And what's even more interesting, over time, for some reason, they like to periodically take this information, the allocation table, and modify it in some way. At the moment we support the main file systems found in DVRs, but over the last three years a certain trend has emerged. There seem to be widely used file systems already, but for some reason manufacturers like to change them periodically. Just last week — our program is currently being piloted in several agencies — and from the feedback we received, in the past week alone two new varieties of file systems were identified: one we've classified for ourselves as TSFS, from Tantos DVRs, and we also came across a TESAM recorder, a real hybrid of the VFS file system — it uses a structure from VFS, VFS2 — and right away, let's say, the program couldn't recognize it, but after running experiments on our bench, the video became accessible, and we'll add this system in the next release.
On top of that, our programs are the only ones that can work with Dozor video recorders, which are quite widespread in various law enforcement systems. Patrol police use them, the Penitentiary Service (FSIN), etc. — you've seen them. We can work with those too and extract information from them.
Wrong direction again.
Briefly about the functionality that's already been implemented. We can work both directly with the storage media — hard disks — and with previously acquired images of storage media. We can work with data that's present explicitly, and we also do carving. The only thing is, in the course of our work we realized one thing: the method we've implemented, the carving, is very deep and takes quite a long time. So we recommend from the start: first we scan the disk, grab the data that's present explicitly, and only then start searching for deleted data. Plus, soon we'll be adding — in test form, it's already passed testing — simply pulling out parts of the video stream by signature search. This has proven itself well precisely on recorders whose disks had been initialized.
In testing we tried initialized disks; when we run them through carving, the process still takes quite a long time — a one-terabyte disk takes us about three weeks to process, unfortunately.
The file system is built in such a way that we can't speed up the process yet, but that's exactly why we decided to add another algorithm: a search by video stream signature. We've tried it, it works quite well, the main data you need is already accessible.
After processing the image we get the disk contents, and we immediately sort it by camera, by date, by time; there are various filtering and search options. And we tried to make the interface, the program, accessible and easy to understand not only for specialists who deeply understand the structure of the process — forensic experts, for example — but we tried to make it intuitively simple so that even untrained officers could, in principle, analyze the disk contents and conduct an inspection. For example, so that in certain cases an investigator could, with no specialist, conduct an inspection and quickly review the disk contents. Because, as you know, in practice specialists are in high demand, and you can wait a long time until one frees up. And, well, since it's perfectly safe, you can look at the contents yourself.
So there's a video viewer, you can view the contents right away, export what you need, and we can export video files individually, and we can also do a frame-by-frame breakdown of a specified region, of the frames you need, for further analysis or for attaching to the report. Plus we also generate a report, quite a detailed one, with information about the disk contents. But there's also a mechanism for deeper examination of the disk contents: there's a built-in hex viewer, which you can use to analyze part of the information and obtain any additional information you need, in case for some reason it wasn't found by our automatic mode. In such cases, please let us know as well.
And briefly, let's go over once more what we've already done to make our program work. We reworked the user interface, simplified it a bit taking into account the feedback from users who had already been using our program. We updated and added a logging mechanism for all the actions the program performs. We optimized the file system detection algorithm. And now there's additional functionality where we can manually specify right away which file system it is. We can even look manually in the hex editor at what's there, if auto-detection didn't work. We look at the disk contents, and if we see clear signs that a file system is present — say, just last week we came across a disk that had been initialized perfectly normally in the recorder, but we noticed an interesting peculiarity. Before the disk was put into the recorder, it had been used as a storage drive in some personal computer, after which the user pulled it out of the PC and moved it into the video recording device, the DVR.
Then it was formatted using the recorder's own tools, it recorded normally, but it turned out that when the recorder was creating its file system, seeing that sector zero was already occupied, it initialized it not from sector zero, as it usually does, but a bit further on. Starting somewhere around sector 40, its initialization table began, and in automatic mode we missed it. But we tried the mode where we manually say, yes, this is exactly this file system, and the program ran and picked up the file system, and everything worked normally from there. We added support for working with E01 images, because, well, everyone knows it's a widely used forensic data storage format.
We optimized the algorithm — originally, when the program first came out, ordinary file systems like FAT32, that is, ordinary car dash cams, basically weren't recognized. Well, the program's original concept didn't call for it, but one of our users asked us — they wanted the program to handle ordinary file systems too, and we added that mechanism. And later I'll say what else, basically, what for, and what exactly we now plan to finish implementing. We already have test builds ready; I hope we'll include it in a release soon. And most importantly, we've developed a new module for analyzing video data that uses artificial intelligence algorithms.
If you've already visited our booth today, you may have already seen it — basically, what results this module can offer us at the moment.
What's special about the module is that we don't tie it specifically to data obtained from VR-Expert; it simply imports ready video streams directly. And those video streams don't necessarily have to be obtained with VR-Expert — they can come from any storage media you have. Say you've analyzed a phone, and in the phone you saw there are some video images — and not just video, actually: our module works not only with video, it also works with ordinary static images, JPEG, BMP and all the formats. And with our new module we can analyze those too. That is, we can either load a single file into it, or we can load a batch of them and run the analysis. And the result of the analysis — moreover, we get it as ready-made fragments of specific images, I'll say which ones in a moment, and then we store all the necessary information about the results in an SQLite database, which later lets us run various data analysis and cross-match the data.
So, the first thing it's needed for is detecting different types of objects in an image. Either across the whole image, or we can define separate zones that the processing will work on. Why the separate zones? To improve the results considerably and to solve specific tasks.
Also, at the request of one of the units, we added a feature: if an object enters a monitored zone, and if you enable the tracking function, then once it enters that zone it will continue to be detected everywhere it has been spotted.
So, at the moment we already have the following classes: people, bicycles, cars, motorcycles, buses, trucks, and we've also added a detection feature.
Right at the start of the job we specify whether to search only within the area or to track the object's entire path. And we've also added license plate recognition. And specifically for this object we added the option of applying an unsharp mask filter and magnification. Again, plate recognition depends on the image quality. And now we're also adding a function where, once a plate is recognized, we accumulate the license plate images from that object, so that later we can pass that data on to the VD-Expert program, which is designed for video-technical forensic examinations and research, and which also includes image enhancement analysis functionality. That's exactly one such method. I think those who've worked with it know this mechanism was widely… You can do it manually in Photoshop, of course, or in GIMP. This mechanism is well implemented.
It used to be in the foreign Amped FIVE. Now this mechanism is implemented quite successfully in VD-Expert. I think if there are people here working in the video-technical field, familiar with it, I think you know this program, you're a forensic expert, a good enough professional.
So in this case we can determine the plate's contents with the help of a video-technical expert. For now our model is tuned specifically to Russian license plates. Going forward we'll be adding models for other countries.
What happens after the processing is finished? We've set the parameters we need for the processing. After the program has finished its run, we can now export, either by class or by a specific object, the resulting image and the information stored in the database. We can save only the individual frames we're specifically interested in. It all depends on what you want.
So what is our program's mechanism based on? We use modern PyTorch models, and moreover, we create our own models. Sometimes, at your request, we can add some specific category you're interested in. Also, during analysis, right away, if you already have some specific model you're interested in, you can load that model during analysis, and the program will work with precisely that model, the one you loaded yourself. We think it's a fairly convenient solution. If you have any suggestions or wishes on this topic, please also come up and talk to us. We'll gladly take your opinion into account and try to implement it.
We added GPU support for data processing, which let us significantly increase image processing speed. In our tests, when working with a single object, a 30-minute video initially on the CPU alone, the processing took, unfortunately, 25 minutes, but when using a graphics card, one object, a 30-minute video, takes around 3 minutes. And there we already have our result. Again, you have to understand that it all depends on the input image quality. There are no miracles: if little information was captured in the video, unfortunately, there's nothing we can do. That was what we've already implemented. And now, what we plan to do in the near future and our plans going forward.
First, we're currently training additional models. In the near future we'll add electric vehicles and certain types of weapons. If you have any other wishes or ideas, we'll also gladly hear your opinion and implement it. There will also be additional filtering by object color — car color, clothing color. That will also be an additional classifying feature that we'll take into account in searches. Face search functionality will be added soon, and much more. Here, of course, we could go on forever inventing scenarios to our own taste, but the better approach is your feedback — which tasks you encounter most often. That's what we'll implement. We're waiting for your feedback.
Also, in the near future there'll be a full port of the engine to a cross-platform system. On Linux, the program started working in test builds roughly six months ago, but, let's say, the engine wasn't quite, let's say, fully ported. We ran it through workarounds; now it'll be fully cross-platform. That is, the engine will work fully both under Linux and under Windows. And we're testing on all the operating systems recommended for use in the internal affairs agencies. And I think we'll fully implement that by the end of the year.
We also plan in the near future to finish what I mentioned earlier about car dash cams. What's the most painful problem with car dash cams? 90% of car dash cam tasks are dash cams that come in for examination after road traffic accidents with the last file's recording unfinished. I don't think it makes sense to explain the recording algorithm. I think you all know why this happens anyway. First comes the video stream, and only then the metadata about where that video stream is located. So, manually, we've already almost fully implemented this algorithm for most file systems. Now we're finishing automating this process, so that the program recovers this data practically in automatic mode. But if you happen to run into such a task right now and can't solve it, we have our own lab. Please get in touch, we'll help you obtain the video information you need.
And now, briefly, let's go over where our program matches its foreign counterparts and where it differs. Well, support — extraction from various data sources — here we match. On most of the points we analyzed, we match. Except that we found the Korean program doesn't always work correctly with password-locked DVRs. But the problem is solvable. We know how to bypass that, and we do. This mostly applies to the "Dozor" line of DVRs. We can crack and bypass passwords on "Dozor".
MD-VIDEO has no explicit support for proprietary formats; we have it, and DVR does.
The version of DVR Examiner available in Russia has no artificial intelligence capability; we have it, and MD-VIDEO has it. Recovery of deleted video files is an implemented feature in all the software products.
Here there's parity too. We tested the speed of our program against the foreign counterparts. In terms of speed we basically match them. Unfortunately, deleted files are extracted slowly everywhere, but to speed things up we're now adding signature search as well, and on most DVRs, on most video streams, that'll be useful, it'll work.
And we also currently have video storyboarding in development, which isn't explicitly implemented in DVR Examiner or MD-VIDEO. A HEX viewer is also implemented only in ours. And we're also now working on support for, let's say, RAID systems found in DVRs.
Unfortunately, those are showing up more and more often these days. And most importantly, only our program has a Russian-language interface. The foreign counterparts don't have Russian in the interface at the moment.
And after all, only our software is registered in the Russian software registry, but we noticed that for full-fledged operation you still need fairly high-performance hardware.
On weak machines it runs quite slowly. We're now fully defining the system requirements that will be the most future-proof.
And as is our tradition, given that, for the most efficient operation, we generally prefer the software to ship together with hardware optimized for working with our software, which will let you get full access in any situation to the necessary information and, where possible, get it fairly quickly. The kit includes, where possible, various disk imagers and write blockers to ensure data integrity. Quite often. What are the imagers for? In the kit we deliberately prefer to supply additional storage media. Why? Quite often, from my own experience I can say, it has happened: you absolutely need the information from a specific DVR. Large volume, requiring lengthy analysis.
And there it is, a DVR from a garage cooperative. The garage cooperative had to be left without surveillance for a week, because at the time it took a week to pull the video. If there'd been a VR, it'd have been much faster. So, even arriving at the scene, you'll have somewhere to copy the information and then examine it later at your leisure.
We also include a set of tools in the kit — for all occasions. And if needed, the hardware kit can also include our software. If the drives in the DVRs are already quite problematic, the kit can also include ACE Lab equipment for working with worn-out DVRs. In an alternative configuration you can also use foreign software, for example, from SalvationData or from GMDSOFT, MD-VIDEO — Hancom is now called GMDSOFT. Because the wider the functionality and the more software products you have at your disposal, the greater the likelihood that you'll fully cover all the different formats and peculiarities of video data layout, and get the information you need sooner.
I'd also like to say that this year our company obtained a license to conduct educational activities and we've developed training courses on working with software, covering all the software products available on the market for extracting information from various sources, including phones, hard drives, and DVRs. In some cases we can develop additional training programs to your wishes and needs, on some specific topics. If you need them, if you want to look into them in more depth. If you have such requests, come to our booth and we'll discuss it. Or you can write to our contacts and we'll discuss it.
We're customer-oriented and will try to find the best deal for you. And moreover, as I've already said, our company has its own laboratory, where we also help law enforcement officers obtain information from various devices. If you've got some object where you can't get access to the information with your own tools, contact us, we'll try to help you. Always glad to see you.
For any questions, you can contact our Telegram support channel. All the contacts are at our booth; come by, we'll be glad to see you. What questions do you have, dear colleagues? Sergey, thank you. Colleagues, questions?
Alexey, please come over with the microphone.
Hello, thank you for the talk. Could you tell me, there's one point that's still a bit unclear. The AI-based video data analysis module. You said any data source can be added. Do you mean any data source at all? I mean, any video recording, not necessarily obtained from a DVR. You have a set of video files, you can load it, and it'll be analyzed. So do I understand correctly, if, say, it's a file system extraction from a mobile phone, you can also add the archive?
Yes. The only thing is, for now it doesn't go into the archive itself; you just need to pull the video file out of it, hand it to the program, and it'll process it. Thank you. Plus, in the near future an image categorization module will be implemented as well. If you have wishes about which types of objects you'd like to see, get in touch, we'll implement them. There's already a draft version, almost ready. —
— Colleagues, more questions? Alexey, I see another hand over there.
Tell me, please, Sergey, isn't LAN PROJECT working on determining the speed of a moving object? —
— No, we're not, and we don't plan to. That task has already been implemented quite well by the company OT-Kontakt in their software product DTP-Expert.
We work in very close cooperation with OT-Kontakt, and basically it's possible to supply VR-Expert and DTP-Expert as a bundle, which would let you solve this task in one package right away.
Sergey, good afternoon. Could you tell me, can you search by the objects your software has detected?
— Search by object type? It sorts it all by object type precisely, and then we can search by them. We deliberately made it so all the information we've extracted from there is also stored in SQL, and then with SQL queries you can quite quickly find information by a given criterion.
Does it work with additional sensors, for example, a gyroscope, G-sensor? Not yet, nothing implemented. We could think about it, we'll see. One more question. Is there GPS support and visualization, if we're talking about… Not yet. In principle, if it's in the video stream, yes, we could try, we could extract it as well. I think that's a good idea. But that'd be in some separate window, right? Yes, plus we haven't fully finished it. We'll finish it by file creation times, and we also plan to add that when there's a date and time shown on screen, we'll additionally recognize it and compare. We're planning that now too, to determine the recording date and time more exactly. Okay, thank you. —
— So, colleagues, let's send Sergey off with applause. We're running a bit over time again. And colleagues, once again, a reminder about the extra prize draw in honor of our company's 25th anniversary. Come by our booth. —
— All right, thanks.
8. Natalia Kotova (Forensic Science Centre of the Yaroslavl Regional Police, EKC UMVD) — “SpyNote in action: how a mobile spyware trojan is created, deployed and examined”
Scheduled 15:35–16:05.
Moderator's introduction
So, those of you who come to our events regularly, including the events we hold in the winter, that is, "Digital Forensics 2025", which we held this year. And you all surely remember that back then, in the roast zone, we had quite a heated discussion. And actually, the intro I'd written for the next talk was completely different. But today, before your eyes, on stage, young talent: Natalia Kotova from the EKC of the Yaroslavl Region. Let's welcome her with a big round of applause.
She'll present a topic that's on everyone's lips today. SpyNote. So, as I promised, here's the clicker. Flip through like this. The bottom one's for slides. Here's the mic. There you go, please begin.
Talk and Q&A
Good afternoon, all. Thank you for coming. I'd also like to welcome those watching the online stream. Today is my first time on stage, and it's a great honor for me to speak at an event like this in front of specialists of this caliber. Colleagues, the talk turned out quite packed, so those in the back rows, don't be shy, come sit closer, it'll be interesting.
So, first, let me introduce myself. I'm an expert in computer and radio-engineering forensics and research at the EKC MVD of Russia for the Yaroslavl Region. A year ago I graduated from the Ministry's university with a specialization in computer forensic examination, so I've been performing them for a year now. When I started work, the first thing I ran into was a flood of examinations related to phone fraud. I don't think I need to say anything about how pressing this problem is. And it just so happened that what came up for examination most often was SpyNote. Briefly, what it is.
SpyNote is the name of a threat family. You can see the full antivirus verdict on the slide. It's malicious software of the RAT trojan class for devices on the Android platform. According to Kaspersky's reports, for several years now SpyNote has been in the top 10 verdicts and its popularity is growing steadily. As you can see, it's been around a long time. In 2022, the source code of its builder leaked online, of one of the versions, which only increased its popularity.
Most often, SpyNote is used to intercept SMS messages with confirmation codes. The cover stories vary. Doctor's appointment booking, an antivirus, a parcel tracker. On the slide you can see samples of the apps that were found on the devices under examination. Only the icons and names change. At the same time, SpyNote has a fairly broad set of features. I suggest we take a look at its capabilities in action.
Without any VPN or SMS registration, online you can download the SpyNote builder, version 6.4. It's a fairly old version, possibly the very one that leaked. The APK files it builds differ from the ones seen in the wild, but they are fully functional. So, you can create your own trojan in a few steps. First, pick an icon, a name, and a version. Second, configure the host address and port. Third, select additional properties and build the APK file. All of this happens automatically. Binding to another APK is also available, in which case the resources of the other app are included in the package.
After the app is installed on the device and granted all permissions, the infected device shows up in the admin panel. In the drop-down list, you can select the available modules. For brevity, I'll show only the most interesting ones. The first three modules on the slide are for viewing SMS message contents. By the way, I want to note that sending SMS is not supported in this version.
Also viewing the call history and viewing details of contacts and phone numbers. Next we see a file manager that lets you interact with the file system without root access. Here you can copy, delete, create, and add new files. There's also covert access to the cameras. It's invisible to the user, and it works even with the screen locked.
Another interesting feature is the keylogger. It involves logging every event of the user's interaction with the system shell. For example, opening and closing apps, navigating the system, notifications, including system ones, keyboard input contents, and so on. All of this not only expands the possibilities for user actions, for tracking the user's actions, but also makes it possible to intercept sensitive information.
The app also has access to some device settings. For example, you can lock the screen remotely, do a factory reset, and even set a lock screen password. All of this without any user confirmation. I tested it myself.
Other features include access to geolocation, the ability to make phone calls, microphone access, a chat with the user, a terminal, and access to account information.
Now let's talk about doing forensic examinations. The typical questions for such examinations are the presence of files detected by antivirus software. And nothing else, no "malicious viruses". That's the only correct wording of the question.
They're also interested in the presence of remote access programs, call and SMS history, message history via messengers and social networks, and web browsing history.
At the extraction stage, first of all we need to obtain either a physical image or a full file system. Next, we scan the extracted data with antivirus software, using at least two different tools.
If the scan detects an APK file, we perform static and dynamic analysis. And, of course, we review the extracted data. Here I want to draw attention to the importance of studying the timeline.
For yourself, you still need to establish the fullest picture of what happened. So we look at and compare, cross-reference all the facts: chats, calls, system artifacts. The most convenient place to do that is the timeline. Now I want to go into more detail on static and dynamic analysis.
The main tasks of basic static analysis include extracting and decompiling the APK file, studying the metadata from the AndroidManifest file, namely the package name, version, minimum and target Android versions, analyzing the app's structure based on the manifest file, which includes analyzing activities, services, and receivers, identifying protection methods such as obfuscation and packing, analyzing the source code to identify potential functionality, and analyzing string resources to look for, say, network addresses.
The tasks of basic dynamic analysis are, first of all, installing and running in an isolated environment, monitoring the app's behavior: what permissions it requests, how it interacts with the user, what its interface looks like, how it works under different conditions, for example, with and without internet access.
Network activity analysis is a must, its main goal being to identify the network addresses the app communicates with. And also file activity analysis, to see what files the app creates or modifies.
There was already a question about tools today. For the tasks of static and dynamic analysis, a wide range of tools is available. I think many are familiar to you. To decompile APK files, you can use basic utilities such as apktool, JADX, dex2jar.
For example, the GUI version of JADX lets you view the decompiled code and the manifest file. In that case, the analysis rests entirely on the forensic expert. To detect obfuscation methods, APKiD is used, and strings, grep, and other tools are used for searching through strings.
For dynamic analysis, first of all we need a virtual environment. There are also plenty of options here. And for network traffic analysis, the choice of tool depends primarily on the network protocols being used. Mostly it's HTTP protocols. In that case a proxy is enough, for example Burp Suite or ZAP. For analyzing TCP and all other packets, Wireshark is of course a better fit.
I also can't help but share a find of mine. This is Mobile Security Framework. It's open-source software designed for assessing application security. However, it works quite well for our tasks too.
It provides automated static analysis and semi-automated dynamic analysis. The tool isn't universal, but quite convenient for initial analysis, and sometimes it does the job alone. It has a graphical interface, which has become familiar and convenient by now. It's deployed as a Docker container.
And it additionally requires an emulator for dynamic analysis. If interested, the link on the slide lets you try its online version for static analysis.
Let's look at static analysis in more detail. The main source of information about an app is, of course, the manifest file. Its specification can be found on the official site for Android developers, link at the bottom of the slide.
From the root manifest tag we get the app's package name and its version. Next we look at the set of permissions. For example, for the APK we built, we see that the app needs access to contacts, phone, SMS, geolocation. The list on the slide is incomplete. Next the manifest describes the structure. An app must have at least one activity.
The intent-filter of this activity, for example MAIN and LAUNCHER, indicates that this is the main activity, accessible via the launcher shortcut, exactly when we tap and the app opens.
The app can also have background services, which are described in the service components. For example, our app has a service that is declared as an accessibility service. That means the app can be activated as an accessibility service, and that will let it access, for example, the contents of the user's screen.
The app can also have receivers that react to system broadcast events. For example, monitoring incoming SMS and monitoring device boot. But you should understand that all this only indicates that such events are being tracked. To establish exactly what they're used for, and whether they're used at all, you need to look at the code inside the declared classes.
Static analysis is convenient to do in Mobile Security Framework. It gives us a fairly extensive report. I'll run through the most essential parts. First comes information about the file and the app, including all its hashes, display name, package name, main activity name, app version. You can also view the decompiled manifest, view the decompiled and even download the decompiled Java source code.
Next is the list of permissions with their descriptions and, importantly, with links to the source code where these permissions are used.
Next comes the list of Android APIs used. A quick explanation here. The Android API is a set of interfaces provided by the Android system for accessing its functions. Comparing these API calls with, say, the permissions requested in the manifest lets us assess how much these permissions are actually used in the code.
Another interesting part is malware analysis. It includes APKiD analysis and behavioral analysis. APKiD analysis reveals the app's protection mechanisms, for example against running in a virtual environment, as well as which tools were used to build the app.
Behavioral analysis consists of matching typical code patterns against a rule base and determining the app's potential behavior on that basis. As you can see on the slide, the report lists such matches. They have short labels you can use to orient yourself. And again, links to the source code, with the function highlighted that's responsible for that functionality.
The report also includes the results of searching the source code for URLs. But keep in mind that it searches specifically by regular expression. Even if the URL is split in the code, domain separately, protocol separately, it won't find it anymore, but sometimes you get lucky. And it will find a complete URL, if there is one.
I'd count it as a drawback that it doesn't search for IP addresses. That would be more convenient than searching for them manually. We'll come back to that.
And the report also lists right away the strings found in resources and code. That can be useful, but I wouldn't say it's convenient to look at them there.
Let me dwell separately on searching for network addresses. It's quite a relevant task.
The main sources for the search are the app's resources and the source code. We can search strings by keywords, for example host, port and others. IP addresses can be searched for by regular expression.
If a direct search brings no results, you can start searching the source code by analyzing the standard Java networking methods. Here again, behavioral analysis will help us find them. Also, critically important information is often encrypted or encoded. In that case the code will contain methods responsible for decrypting such values.
SpyNote performs its functions exclusively on command from the C2 server. Also, when building the APK file, if you remember, we specified the host and port values. Consequently, the app must contain these values somewhere. And in our sample this information is stored in the resource strings. Here, you can find it here.
Let's move on to dynamic analysis. In SpyNote's case I used a combination of Mobile Security Framework, the Android Studio emulator with Android 10, and Wireshark, since SpyNote uses TCP packets to communicate with its server.
There's no other way to see it.
After installation, the app immediately asks for confirmation in the accessibility menu. Then it redirects you to activating the app as a device administrator. It also creates a folder in the user directory where it stores the APK file we used for binding. And, interestingly, most importantly, the SpyNote log file.
If we look at this log file, it contains both text data and Base64-encoded blocks. It writes all its keylogger data there.
And by decoding these blocks, for example, we can get app icons. If you test the app without internet access, which is the forensically correct way, you'll only see DNS queries for the C2 server's domain name. That's already enough to discover the address and confirm the results of the static analysis. But with internet it gets more interesting. First our domain, of course, resolves to an IP address, and then a TCP connection is established with it. The data is transmitted in compressed form.
If you look at the packet payload, by default a certain set of data is sent that's needed for display in the admin panel. The slide shows an example of how an image file can be extracted from the packet data. You can see the thumbnail here.
And further interaction happens exclusively on requests from the server. And for dessert, a bit of forensics. If the app is installed, it's all simple. The data directory will be located in /data/data. Information about the installed app, the installation date and time, its display name, we can get from the database files frosting.db and verify_apps.db. These databases are basically always there. And, of course, the main file containing information about all installed apps is packages.xml.
Among other things, it also lists the install initiator and the installer app.
It gets more interesting if the app is not installed. In that case, searching by the app's package name comes to the rescue. I generally recommend always doing it, since a lot of interesting artifacts may turn up. And they can differ from one mobile device manufacturer to another.
For example, on Samsung devices you can find info on an app's installation and removal in the battery usage log.
The Android package manager log contains even more information, including, again, who initiated the installation.
Well, what do you do if neither the file nor the app is on the device anymore? Here, not an advertisement, help can come from a built-in antivirus, for example Sberbank's app has one. The file 30.db records information about the detection of the app file, with the verdict, package name and, accordingly, the time. So we can prove a malicious app's presence without it being present. That is, that it was there.
If SpyNote was run, its log file will most likely be present on the device. For example, on the slide you can see a fragment of its decoded contents.
That brings my talk to an end. The topic is quite broad. I tried to cover all the key points. What came out is a small practical guide that you can use during an examination. Thank you for your attention. I'm ready to answer your questions and discuss everything.
Great! —
— Most modern malware uses crypters and code obfuscators. As I understand it, the source code here is open, there's nothing like that. But have you faced this in your practice, and how do you deal with it? Also, a question roughly on cryptography. If we have a pinned certificate, SSL pinning, did you do SSL unpinning, anything like that, to read the traffic?
— Thanks for the question. Regarding obfuscation. Yes, in SpyNote it's used quite heavily, but here it's really about how much you need to dig into the source code. Whether you have the skills for it, whether you have the time for it and how much you need to dig into the source code. Mostly the obfuscation shows up as enormous, strange, meaningless names for all the classes and variables. So analyzing it manually is quite difficult. That's why I found it convenient to use an automated solution, for example one with behavioral analysis. There you can quickly see where a network socket opens, where the SMS handling is. Again, as much as it's needed. Mostly the task boils down to finding the IP address of the C2 server.
And about... repeat the second question? The second is about traffic encryption. If the app uses its own certificate, it encrypts traffic with it, and a substituted one, it won't respond to it, even if we load a Burp CA cert, say, onto the phone, we still won't see the decrypted traffic. Yes, with traffic, look, the thing is, since this type of software is client-server. For example, when conducting forensic examinations we try not to let the samples under examination out onto the internet, and so for the most part we don't really face the task of traffic analysis as such, because, let's say, the probability of getting a response back from the C2 server after some time has passed is, well, fairly low.
There have been cases, of course, when the scammers used that chat and wrote to the forensic expert. But usually that's not the task at hand. I gave this example here because I had both the client side of the app and the server side. I could trace that traffic. In this case no encryption is used. Only Gzip compression is used, which is fairly easy to decode to get the contents of the packets. It's just that some apps won't give you the C2 server address if they realize they're offline. They'll make a request to, say, some server API to determine what their external IP address is. If they don't see a response, they won't connect anywhere. Yes, that's a very good question too.
In that case there's another option, to avoid letting it onto a real network. If the app simply checks for a network connection, you can, for example, use, say, a pair of virtual machines, where one virtual machine runs a network emulation. That is, for example, the INetSim software lets you emulate a setup where any request will get a positive response from a server with stubs. Then we can pass that kind of check. —
— Natalia, thank you very much. That was great. Let's give Natalia another round of applause. We're just pressed for time, but we'll still have the roast part with you.
9. Andrey Shavlovsky (Forensic Expert Centre of the Investigative Committee of Russia, SEC SK) — “Examining information by dynamic analysis on macOS-based personal computers”
Scheduled 16:10–16:40.
Moderator's introduction
So, actually this ninth conference has turned out, actually, a bit strange for us. Because, as you may have noticed, through pretty much the whole talks section we, basically, this year, didn't mention Apple devices much. Maybe that happened because at the last conference my colleague said they're impregnable. Maybe for some other reasons, but back to their impregnability, and we'll bust that myth right now. Please welcome Andrey Shavlovsky, Investigative Committee. Let's support him with applause and talk specifically about Apple and macOS. Andrey, the microphone and the clicker.
Talk and Q&A
Good afternoon, dear colleagues. I'd like to thank the organizers for the opportunity to speak at this conference. My name is Andrey Shavlovsky, I'm a digital forensics specialist, I've been doing computer forensic examinations for 8 years now, digital forensics. And in this talk I'd like to cover some issues related to the method of dynamic examination, memory analysis on personal computers running Apple macOS. The topic is quite relevant. Actually, I've touched on some of the issues related to it. You could actually talk about this for quite a long time. But overall we'll talk, well, basically digital forensics has well-known, generally accepted approaches that don't need extra interpretation to be discussed in detail.
That is, the object is examined, the integrity and immutability of the data it contains is ensured, an exact sector-by-sector copy is created, then the data it contains is analyzed, the data is interpreted, and the results obtained are saved. That is, overall there are plenty of methodological guidelines on this. Overall this is, roughly speaking, the basic static analysis that is performed on disk images, on electronic storage media.
Overall, the previous speakers described dynamic analysis in detail, and static analysis. Overall it's clear that static analysis is the examination of information on a computer system or storage medium that isn't running, when it isn't launched for execution but is examined as is. And dynamic analysis is when a running object is examined, when it's launched for execution. That is, there are some indicators of malicious behavior and others. Overall, static analysis is the gold standard of digital forensics. Most examinations, of course, are done precisely by static analysis, because as a rule you can ensure integrity and immutability of the information, there's repeatability, there's the possibility of repeating the examination.
The results, accordingly, are reproducible and repeatable. But there are also problems, when some data may be missed from protected memory areas. So, for example, overall, dynamic examination methods are applied to the Microsoft Windows operating system, when the user's password hash is pulled out of the registry files, it's cracked, that is, it has fairly weak cryptographic strength, and then an exact copy is either virtualized or examined by other means. Data is extracted from browsers, authentication tokens and other data that may have forensic value.
So what about macOS? Overall, what's known about macOS is that it's a proprietary operating system with its own proprietary file system.
It has a very important built-in mechanism for storing keys and passwords, the Keychain, access to which is possible only after authentication with a password or biometrics. An analogue of Windows DPAPI.
The directory structure is different, somewhat similar to Linux and to Windows. Overall, the file structure has a system level, a user level, an application level, and a level where temporary and cache files are stored.
By analogy with Windows, when a user is registered a home directory is created, conventionally called the user's home directory, whose name matches the user name. By default it's located at paths similar to Windows, but it can, for example, be moved to external storage as well.
Data, application settings, configuration files, configuration data, accounts, activity logs in macOS are stored in plists. So that's no secret. But plist files come in different formats.
The main formats are XML, a text format with tags and indentation. It's used, accordingly, for editing and debugging. And the more popular format is binary. This is, essentially, a machine-oriented format. It's not readable without special conversion.
There's also the JSON format, but it's used less often. So, if you open in text editors a binary plist file and a text one, the XML format, then it's clear you need to convert from the binary format to XML. Why aren't all plist files in XML format? Well, there are two main reasons. It's, of course, that the binary data format is more compact, smaller in size than XML, doesn't require indentation, tags, like XML does. And the processing speed of binary plists is much higher than XML, so by default plist files are used precisely in the binary format.
So, if we move on to the question of obtaining the user's password, that is, where its hash is stored, then in general the user account data is stored in one of the system directories, namely the dslocal directory.
[music]
So, when a user account is registered, a separate file is created for each registered user, whose name matches the username. This file is in binary format, and accordingly, to view it, you need to convert it to XML format. Accordingly, we're interested in the ShadowHashData data section.
Once access to the file is secured, and the conversion from binary to XML is performed, you can actually work with the data we see. To convert from binary to XML format, you can use various commands. For example, the plutil command, a utility built into macOS, in the bash terminal, which you can use. You can use special software, text editors with plugins.
After the conversion is done, ShadowHashData stores the data, that is, it's encoded in Base64. It also needs to be converted.
Also, in principle, these are accessible methods. There are many online decoders, there are software tools that perform decoding. This can also be done in the terminal. And as a result of decoding we obtain information about the password hash. Accordingly, we see that it's encoded using the PBKDF2-SHA512 hash function. This is a cryptographic algorithm, that uses the PBKDF2 function in combination with the SHA-512 function. So, a cryptographically strong algorithm. And the data is divided into several, roughly speaking, values. The first is the entropy. This is the hash function's final result. Stored in hexadecimal form.
Iterations. This is the set number of iterations. how many times the function, the hash function, is executed. And the salt is a unique sequence of bytes, which is added to the user's passwords in order to then perform the hashing.
What's the distinctive feature compared to, for example, Windows? Windows uses the NT hash, which is stored in registry files. Accordingly, it's used without salt and without iterations.
And besides that, the NT hash is computed with the MD4 hash function, a fairly outdated one, its cryptographic strength is very low, so, the password brute force speed for NT hashes is very high compared to Macs. The salt prevents, for example, the use of rainbow tables. That is, for example, if two different users have the same password in macOS, then the hashes of these passwords differ, because the salt is always unique, unlike in Windows. That is, for example, any leaks that are used to guess passwords, rainbow tables, they're very often used to guess passwords for Microsoft Windows accounts.
Accordingly, the passwords, the hashes stored in macOS, are more cryptographically strong. And the brute force takes longer.
Also specifically in hexadecimal form. And the final result, of course, is the entropy. That is, after performing the required number of operations, this entropy should come out. If it matches, then the password fits.
We've covered the step-by-step extraction and conversion of the password hash and preparing it for further brute forcing. So what tools can you use to automate this process? You can use special scripts, which are fed the corresponding plist file as input with the ShadowHashData, that is, specifically for this user account. or, accordingly, get as a result a format, a representation in a format that you can then brute force in hashcat. In principle, for the format we can also use software tools. In Mobile Criminalist Expert I didn't find that it can convert the password hash data for a macOS user, but maybe they've fixed it by now.
I used other software tools that are available, but overall I ran into a problem — we won't name this software tool, but I ran into a problem where the software tool in question converts the data, but displays it incompletely. That is, part of the hash is lost, it's missing. Accordingly, it's not possible to brute force such a hash. In general this situation clearly shows that automated software tools don't always give an absolute result. You have to be able to double-check this data.
So, the data is obtained, hashcat is launched, it's passed the corresponding file that contains our password hash. It's given the information about which hashing algorithm to run this brute force with. In this case it's 7100.
In the instructions for hashcat, all this information is available, it's publicly accessible. And overall the guessing is done either by mask or by dictionary.
And you can get a result, if, of course, you manage to gain access. Not everyone likes working in console programs. You can also use graphical interfaces. There are programs built on hashcat with graphical interfaces. One of them is MK Brute Force. An excellent program, it's developing, a lot is being added. Unfortunately, at the moment it doesn't support the 7100 password guessing mode. And to test it for guessing the password of a user hash in macOS wasn't possible. But let's hope that later there will be updates and such a capability appears. For this, other tools were used that are also built on hashcat and can perform user password guessing.
Accordingly, our file is passed as the input parameters, which contains the password hash. The user selects the required mode. So, the mask, or the dictionary, to run the brute force by. And, accordingly, you can already clearly see the guessing speed, the graphics card temperature, and so on. Of course it's best to brute force not on a single graphics card, but to have a GPU cluster overall. In principle, here you need to use more advanced methods too, in the sense that simple brute force is fairly long, labor-intensive, and the time can run into years, which nobody needs, so, of course, one compiles, well, basically, you need to build custom dictionaries, Based on data from notebooks, data from digital devices, I can recommend the web resource wikpass.com, it contains many custom dictionaries that can also be used for brute forcing.
In the end the password is established, and then you can move on directly to the examination itself. The two main methods are: one where the object is virtualized, that is, its virtual environment is set up on a previously taken exact copy.
If, of course, creating a virtual environment, a virtual copy of the object, was not possible, then, at least, the law permits examining the object with changes made to it. But at the same time, clearly, if we're talking about a forensic examination, then prior permission must be obtained from the initiator of the examination to make such changes. Because there's Article 57, which tells us that a forensic expert is not entitled to use examination methods, that could cause the full or partial destruction of the object, or a change in its main properties and appearance.
Accordingly, having the user's password, you can decrypt data from the keychain, pull out authentication tokens, pull out passwords, that is, obtain far more data than, for example, without the password. And accordingly, in some cases we can directly examine the information content itself, as the user sees it. We can look at, for example, messengers, desktop ones installed in the operating system, and that may not be parsed by static analysis, but that can be extracted dynamically, including with the help of the same Mobile Criminalist Scout.
It's given the appropriate privileges, and you can also extract this data, tokens, passwords from browsers. And this is also quite interesting and pretty good. Overall, you have to understand that automated software tools, they make our work much easier, because the volumes of information are large, and honestly, working through that quickly is quite difficult. But you have to keep in mind that software tools can make mistakes, that is, their code may contain bugs. And their output should not always be treated as absolutely reliable. You need to take a critical view of the results you get, including re-checking them, re-checking the key results, the key information, with alternative methods, including manually, ensuring the verifiability and reliability of the results obtained.
And the methods of static and dynamic analysis, on the whole, they complement each other, they provide a combined approach to examining digital traces, and taken together they let you extract the maximum amount of digital data that an Apple macOS computer may contain. And they help make the examination as complete and comprehensive as possible, by extracting the maximum amount of forensically significant information.
Thank you.
Andrey, thank you for the talk. Questions from colleagues. —
— Yuri Mikhailovich, please. So the disk wasn't encrypted? How did you pull the file out?
In this case it was a Fusion Drive. We managed to reassemble it, and the data was pulled out, extracted. Obviously, if we're talking about encrypted disks, then if you decrypt it, you get the user's password as well. That's all logical. But if, as in this case, there's no encryption, you can try to gain access in various ways. —
— Next question. Good afternoon, thanks for the talk. Here's my question. On Linux, say, we can't get access to the shadow file with passwords under a user account, but on macOS we can? I mean, if we have a live system, say, the password's been entered, we won't be able to read the plist file? —
— Yes, thank you for the question. There are certain difficulties, especially if the device is a modern one, so, there's hardware encryption, the Secure Enclave is used. And here, of course, even with access, well, without root you can't get in, unfortunately. There are certain difficulties. So mainly, of course, if you have full access to the file structure, you can pull out that file, the plist, and then work with it. The whole point is to get that plist file out. If, of course, you can't get it out, then, unfortunately, there are certain limitations here. Well, it's no secret that there are a lot of problems with Apple devices.
And modern iPhones, as we know, are also quite problematic to examine, even if you have the passcode, including extracting the full file system. Well, there are certain limitations, so they have to be taken into account. —
— Right, colleagues, more questions? No more questions. Andrey, yes, we're letting you go. Thank you very much. Let's give him another round of applause.
10. Igor Bederov (T.Hunter, Internet-Rozysk) — “Identifying the owners, administrators and developers of web resources”
Scheduled 16:45–17:05.
Moderator's introduction
You can put the tincture over there, please. And we move on to the final talk for today. After that, only the roast is left. But what's the topic of the talk? Let's muse a little, now that the day is ending. The internet, in itself, is a very interesting place, because sometimes you find things there that you weren't even looking for. Not that long ago, about three years back, a story went around the web that at the bottom of a kindergarten's website someone sold illegal substances. I don't know if you saw it, but there was such a story. So on the first page, kittens, and on the second, a whole drug-dealing network.
Our next speaker, Igor Bederov, will tell us how cases of this kind get unravelled. And he'll do it right now. Igor, the floor is yours. The microphone and the clicker, please.
Talk and Q&A
Thank you very much, everything's working. Right, I've gone back. Yes, once more? Yep, once more. —
— And once more, yes. Thank you all for coming. And indeed, well, from forensics let's try to dive into OSINT. Competitive intelligence: collecting, researching public-source information. It's also very important for us. We've already discussed the prospects of forensics with colleagues in the hallway, especially the prospects, given that, possibly, in the upcoming iPhones, and other mobile phones too, they may drop the USB Type-C port, the charging port, and how forensics is going to develop at all, when there's simply nothing left to plug forensic software and hardware into. So, data analysis is probably going to develop as well. And although at our previous meetings we talked about Telegram users, about Telegram channels, today there seems to be, at least, as I've been told, a request from the audience for research into websites.
The topic seems perfectly simple and obvious. Websites have certain users we're interested in, and our task is to identify the people involved in owning this site, administering it, or developing it. So, the three roles that appear in our investigations are the owner, the administrator and the developer. The one who owns the domain name or the hosting, the one who administers it, communicates with users and posts this or that content, and the one who develops the site's engine, hooks various technologies up to it and administers all of that. Where to start? Let's start with the simplest thing, a pile of useful sources. You can take a photo, or you can not take a photo. This slide will come up again at the end.
My colleagues and I went to the trouble, specifically for everyone who does research and investigations, of creating our own build of the Opera browser, a portable browser that runs from a USB stick and stores your authorized sessions on that same USB stick. So you can take your workstation with you, and in this browser, besides a number of privacy settings, there is also a pile of useful sources, including for researching websites, and for other kinds of research that may be useful to you. Download it via the link, and if someone doesn't like Opera, there's an option to load all the useful sources into any other browser as an HTML file.
So, let's go. A web resource is, first and foremost, a domain. What we see in our browser is its domain name, the one we navigate to. vk.com, gosuslugi.ru, or some other one. All these domain names must be registered. I won't go into the details about the international organizations, about ICANN and so on. A domain name gets registered, and the information about that domain name is stored in various WHOIS services. Everything seems great, wonderful, but naturally we run more and more into domain names being registered extremely sloppily, the owners' details are not verified, not checked, and that poses a certain problem for us.
WHOIS data becomes unreliable, and with the arrival of a thing like GDPR and the restriction of personal data in WHOIS, we've come up against the fact that it says a private person is the domain owner, and what to do with that, we don't fully know either. But when we do OSINT, we understand that all of this can exist in historical retrospect, so when we need to get WHOIS data for old sites, we can dig into the WHOIS data archives, which are kept by a large number of services listed here on the screen. Yes, GDPR has limited the amount of personal data returned in WHOIS, but in WHOIS archives it may have been preserved, and we have to check that data in order to find out the name or the company that owns the domain name.
The next thing is a bit off-topic, but since in August the Russian Supreme Court issued a ruling for business entities, and I think there are probably security service representatives, or future security service representatives, here, obliging commercial entities to detect typosquatting on their own, and online fraud, I can straight away suggest a few simple and obvious resources that let you monitor the appearance of domain names that resemble your organization's domain name. And also monitor the leaks that happen involving your domain. They do it for free. So on one side we have DNSTwister, dnstwist and the like, which let you find domain names with a similar spelling, see if they're active or not, if a site has appeared there, whether email has appeared behind that resource. And IntelX, Have I Been Pwned are services that let you find out if there have been leaks on your domain, and thus protect the organization.
The next thing, after we've talked about the domain name, is hosting, the physical location of our site. Our site is images, texts, some volume of information; it has to physically reside somewhere. It resides on an external server, which is called hosting. To figure out which hosting our site is using, there is also a large number of services, starting with a plain ping, which you can do from the operating system, and ending with a pile of external services.
Now, the most important and central thing, probably, about web hosting, is Cloudflare protection. Originally it was created so that we could minimize DDoS attacks, but in practice Cloudflare is actively used by offenders to hide the actual location of their site. And a frequently asked question is how the Cloudflare protection could possibly be stripped away to understand where our site is physically located. And you can't always strip it away, but to some extent you can, through external services that may have indexed our web resource before the Cloudflare protection was put in place. That's URLScan, that's VirusTotal, that's various leaks of Cloudflare itself that are out there on the internet, that's DNS data analysis, that's analysis of all the other systems and technologies present on our site that verify the site in Yandex, Google and other systems, load new technologies, payment acquiring and the like.
And finally, it's looking for reuse of our site's SSL certificate and favicon. Now about DNS. We have a domain name and we have the physical location of our site. Linking the first to the second is the job of DNS records. I mean, we all remember, some time ago, in 2021, a certain social network, banned and designated terrorist in Russia, suddenly stopped working. That was in October, I think. And journalists were writing to me, screaming, in tears: Igor, tell us, why exactly can you get onto this banned social network while nobody else can? I said I was going straight to the hosting, to the IP address of that social network. So that's what DNS records are responsible for.
Thus, DNS records hold all the information about the servers linked to the site we're examining. Some of them may not be covered by Cloudflare, some of them may be physically located on Russian territory. And we've come across cases where drug-trafficking sites, sites spreading false content, and other illegal resources may have part of their infrastructure based on the territory of our country. We came across quite a lot of that. SSL and favicon, which can also be used; listed here are products that can help us search for reuse of these site elements. And it's not quite a linear story when it comes to trying to identify and find the owners of our resource.
Linked contacts. We often look for contacts in the body of the site itself. And it's important to note here that contacts for a web resource turn up not only on the site itself; they can be in external leaks. First, registration data, archived data, WHOIS. It existed, it's been saved somewhere, it was in numerous scraping and parsing results, and it's available in large quantities on the internet. Second, there's a huge amount of advertising using one site or another, which can also be collected by various crawlers. And finally, the millions-strong leaks that happened here and worldwide; and those leaks can also be linked to domain names, so we'll see the pattern of how the email address is formed, which employee names appearing in those addresses exist on the company's domain, and maybe which passwords are used there. This lets us collect contacts.
On top of everything else, we can guess these contacts using a standard pattern. The domain name plus standard, for example, email addresses: office, contact, admin, support, HR, PR and others. Then check them with an SMTP request to see whether they actually exist.
Web resources also store a large number of external files. These files, of course, can be found using various external services. VirusTotal is your friend here. They can be found using advanced search operators, or dorks, like the ones on the screen right now. And most importantly, these files very often store metadata. One of our investigations, as funny as it may sound for 2025, involved an invasion of privacy. Two people, two businessmen, fell out. One made a website about the other and posted all kinds of scribbles there, nasty pictures featuring his competitor, and he made those pictures on an iPhone, and the iPhone dutifully saved in the metadata the geolocation of where those pictures were made. This was in 2025, as funny as that may sound, and the geolocation was in fact the home address of the person who was the main suspect.
Among other things, metadata, as we know, also stores data about the camera, the time and date of the shot, and other information that matters to us. Hyperlinks are another important part of the website under investigation. Hyperlinks can be both external and internal. Internal ones are links within the site, between individual pages of the website. And some of them may be hidden, not public. In that case we examine the robots.txt and sitemap.xml files to find out what other pages exist on the site that we can't see. They are often of interest to us. And finally, external hyperlinks. They can also be useful to us, because these are links to associated social networks, to external file-sharing services. For basically any file-sharing service, you can determine which email is tied to it, even by OSINT methods.
All this lets us identify the people who administer the site. For example, if we don't know and can't find the site owner's or administrator's details by simple means, but it has an associated group on the VKontakte social network, then using the utterly trivial InfoApp application we can obtain the profile data of the administrators of the group linked on VKontakte, and that will be much faster, more convenient and more probative. The next point, also quite important, is the technologies used on our web resource. What's included? Numerous bank acquiring services, chatbots, feedback forms, analytics counters, advertising ID codes, and so on and so forth. Everything we try to cram into our web resource in order to collect information about the audience, build feedback with it, in order to control it and build targeting — from a marketing standpoint all this works in our favor, but for reconnaissance on domain names and sites it's, of course, a huge minus.
At the very least, we've come across a great many Yandex.Metrica counters placed on sites — on banned sites, on sites spreading false information. And the funniest thing is that all of this can be used for identification. For example, we have a Yandex.Metrica code. First, about 10% of Yandex.Metrica counters are public: you can open one and simply look at the moment it was installed on the site. And when it was being installed on the site, naturally, the only user who came into its field of view was the person installing it on that site. And the second point concerning Yandex.Metrica is that through support — quite wonderfully — Yandex support often, not exactly as a secret, but discloses the email address associated with a given Yandex.Metrica identifier. Just ask them for a hint: "I'm the site administrator, I forgot which email is linked to this Yandex.Metrica counter," and they tell you that email address.
Things like that have happened too. Acquiring. Acquiring services are tied to banks. Then, with an acquiring service installed, you can contact the bank to find out who obtained that technology for placement on the site, and get his login and other registration information. Any site also exists in a certain historical retrospective, so we also turn to web archives to see what it looked like before, what the site's code looked like before, the elements and technologies that were part of that code, what contacts were previously listed on the site, linked social networks, external and internal hyperlinks, and even files. All this will be in the various web archives.
And finally, external traffic. It's often left out of website investigations, but for us the most important point arises here. For example, some sites hosting petitions. You've probably often come across petitions in your investigations calling to overthrow the government, topple presidents, governors and other officials. Here, very often, in about 70–80% of cases, we find that a petition, yes, can be inflated, a petition, yes, is almost always inflated by bots, but as a rule the petition's author initially tries to seed it himself. So we turn to the web resource's external traffic to find out who first posted links to that petition on the net with calls to vote for it. And we often find such people, find them on the social media page where they pushed it, the groups where they posted the calls with that petition, and that lets us find the authors of these very things.
To sum up this whole story, the overall outline of what we talked about today. Very briefly, because the full lecture on website investigation takes us about an hour and a half.
Everyone will get the slides. Useful sources, the bare minimum of what you can use within OSINT to check web resources.
And, going back, once again the link to the browser I showed at the very beginning. It has practically the quintessence of everything we talked about, both in our previous talks and in this one. Investigate cryptocurrency, social networks, websites, Telegram, work in a security department, do forensics — most of the software is generally free, and in this browser you'll find more than 2,000 sources.
Thank you, you've been a wonderful audience.
No doubt about that. So, colleagues, your questions. Igor, you've apparently fired up the audience so much that it's all perfectly clear now. Then let's send Igor off with another round of applause.
[applause]
And we'll move along, bit by bit, toward the very final part of our evening today, namely the roast.
Closing of the online part of day 1
Dmitry Yankovoy (moderator).
But before that, before we move on to it, now is the time to say goodbye to our online viewers. And we'll see you tomorrow, right at the start of our second day. And with you, dear guests, we'll now drift smoothly over to the bar area, taking along some drinks, spirited and otherwise.
[End of the day 1 stream: the "roast", the prize draw and the informal part were not recorded. What follows is the day 2 stream; it starts in the middle of the moderator's introduction to the first talk (with a pause in the recording between them).]
Day 2 — Friday, 12 September 2025: information security day
11. Yuri Barkalov (Forensics Science Centre) — “How does OSINT affect information security?”
Scheduled 11:05–11:35.
Moderator's introduction
And the next topic, which opens our day today, is quite a broad one. It's not Apple devices, it's OSINT, about which lately there's been a lot of interesting stuff in the info space. Interesting not because it's like it used to be — you took something off the net, opened it, found it, and then somehow used that information. It has become interesting because OSINT has now turned into a real gray area that generates a huge amount of debate. And how to use it now is, in fact, not really clear. Let's try, after all, to figure out how OSINT is used in information security, and Yuri Mikhailovich Barkalov will help us. Let's support him with applause, because, Yuri Mikhailovich, we're placing a lot of hope in you today.
Here's the microphone. Thank you. Your clicker. —
Talk and Q&A
Actually, yesterday there was already a talk on OSINT — it was the closing one. So, accordingly, today will be the beginning of the continuation.
So, this is from the prose of life. Okay, down, right? —
— Ah, there it is. Well, a bit about myself. Here again, when it comes to OSINT, there's a bit of information missing. I now also teach at the International Institute of Computer Technologies. On top of everything else, that is. Well, let's move on.
Well, OSINT, as a matter of fact, everyone knows the translation: collecting info from open sources, intelligence. The word "intelligence" is still there in the English name. Personally, "competitive intelligence" or "computer intelligence" is closer to me. Why? Because, well, we do have such a term here, OSINT is international, everyone gets it, everyone brags about it. But I'm not too fond of English-language terms, for the reason that for example, legal proceedings in the Russian Federation are held in Russian. And then explaining what it means in English and in Russian, maybe a young judge would understand.
Well, collecting information from open sources, it's all clear. But the main thing here: comply with Art. 272, and 152.1, 152.2 of the Civil Code, so as not to break the law and not get caught, not incur liability. But actually, I'm not going to touch now on the topics of how OSINT is conducted or what it's for. I want to talk about something else. Actually, well, maybe I decided this for myself, maybe it really is so. But OSINT can be divided into at least two parts. There's professional OSINT. Well, we've got the Positive Technologies folks here, well, all the rest — anyone who does professional analysis of incidents in the field of information security, knows what OSINT is for and how to use it.
And there's civilian OSINT. Any of you, everyone tries to find something online. That is, everyone is a subject of OSINT, both on the side of receiving information and of providing information. And here's what's interesting, if you think about it: did OSINT appear long ago? The first civilian OSINT was probably the grannies who sat — well, the older generation remembers — the grannies who sat at the entrance and handed out information about the persons of low social responsibility living in that building. Well, I hope you understand what I mean. Now those grannies, pardon me, have been replaced by young people who sit on social media, and who knows what they discuss there. And now here's the interesting question, the main thing. So
there are consumers of OSINT, and there are those who supply information there. And since the key word is "intelligence", accordingly, counterintelligence exists. And what is counterintelligence? It's disinformation. So, if it comes to that, look, I'll say a bit more now about civilian OSINT. And by the way, about yesterday's remarks on when I did a forensic examination. So, recently, someone brings me an order appointing a forensic examination, well, a lawyer brings it, and there are questions. I read the questions. Excuse me, Elena Rafailovna Susova — they can take a rest with their question list, which they came up with back then and which some people still use. Who knows what's in there.
Turns out — where did you get the questions from? Alice helped. Alice — that's the Yandex one, right? You see? So it collected something, did something, provided something. And people, well, quite — artificial intelligence or whoever it is, I don't know, but the information is there, it can be used.
Next — that's not all, naturally. Well, it's clear what they collect it for. Well, I put pluses and minuses here, actually. Well, let's take the minuses: preparing attacks. Well, I really did get a call — I mean, a message, then they call. First, a Telegram message on behalf of the head of the Interior Ministry institute, saying: "Yuri Mikhailovich, an FSB representative will contact you soon, you must assist him." Some time later, a call, he calls, obviously asks about something, I say something, and they ask me: "Well, tell us how information security is organized at your institute." I say: "I've been retired three years." "How would I know?" "We know, you tell us anyway." Well, you see, I just collect things like this. I find it interesting to talk with people like that.
Well, we talked, we chatted. Well, obviously, then it comes: "Well, we'll soon invite you to Lubyanka." I say: "Well, fine, I'll come." Olga Vladislavovna told me today: "Don't say what's next, so it won't get out." So I won't tell you what I answered them. Let's move on. Well, it's clear what else they collect it for: blackmail. Well, again, what do they blackmail with, exactly? Well, many of you have seen it: it's mostly, let's say, photos and videos of an intimate nature. Including, maybe, deepfaked ones, I don't know. Such things happen. Recently there was a request to do a forensic examination of two sodomites in cassocks, excuse me.
Was it a deepfake? I had such an examination. But I'll say right away why I didn't do it: because the recording quality left a great, great deal to be desired. Well, the fact remains that this happens too and this is possible. Well, obviously, theft as well. Look, actually, theft of what? Not just theft of information, but theft in the ordinary sense of the word. Look how nicely I'm relaxing at a resort, somewhere far away from home, right? and nobody's at home, so, like, come in, take it, nobody's there, help yourself, and look what nice things, bought this, bought that, so, well, there's something to take. Well, I won't even talk about fraud, that goes without saying, they gather some OSINT, something, well, actually, that'll be discussed later, it's enough for scammers to know some tiny bit, a name, something else, they cite some passport details, which can also be found.
And then what? Then come social engineering methods. How does it go, remember, the film about Buratino? "As long as there are fools in this world, living by deception suits us fine." So what? Here you go: they withdrew the money, it's in your hands, and now put it in a "safe account". Why put money in a safe account when it's already in your hands? Far safer: put it under the bed, under the pillow. But people get brainwashed, well, what can you do, it's human psychology. And by the way, on that note, here's the thing: that information... it's no accident that right now, on the one hand, of course, it's bad and inconvenient, this restriction of foreign messengers, but on the other hand, well, how else do you do it? Because again, people swallow whatever information is shoved at them.
How did the old ladies at the fence, by the entrance, know who lives where, and what the social responsibility of certain residents was? Same here. Who posts the information? Where does it get gathered from? Well, now we can move on to the pluses, for the good things. But I've practically already said that, well, even some of my own students are sitting here, who carry out that first item in certain organizations, well, yes, they collect information, so what. And you: don't post anything bad about yourself online, and make sure nobody else posts it about you. Well, you understand what I mean: well, got drunk, sorry, or something else happened, I'm not talking about last night.
Yes, checking for data leaks, well, obviously, I already said, well, this is probably the most common case of using OSINT from the digital security point of view, well, when you check whether your data is actually online, whether it leaked or not. But crime investigation — here's a very interesting thing, which is sort of where I'm going. Crime investigation I marked with a plus, but you could also give it a minus Look, the flight, now I'm getting to the not-so-good things, MH17, who investigated the incident, when the plane was shot down over those regions. Well, their investigation was done through Google. Actually, they really did have information there, because they assume that people post correct information.
But who, where did they get the information from, who posted it? That's what I want to stress now. Actually, I say OSINT isn't what it was. Or maybe it is, it's still there, the classics exist. I'm saying, about CyberDed there's nothing to even say. They work, they're all great. They have their task, they do their task. There are other organizations. I mean something slightly different. I mean information security, but not that broadly, just from the point of view of the Information Security Doctrine of the Russian Federation.
Why? All the information that's out there online and which, excuse me, you and I as consumers, and everybody else receive, and we receive it constantly. And this information, what can it be about? What can it be preparing? And, excuse me, the Information Security Doctrine clearly says that protecting information includes protecting society from information that is harmful.
That is, involuntarily, every one of us becomes, every citizen of Russia, a consumer of who knows what. So, basically, this information has arrived. Well, you know, here's the first thing one could say right now. A sister-in-law's brother-in-law's nephew said there's a currency reform tomorrow. And he works at Sberbank as an assistant to some janitor. Doesn't matter, people don't think, what matters is the keywords are there. Well, again, that's NLP, that's all social engineering. Create panic, get something done. You'll say, that's not OSINT. And I say, it's possibly not OSINT. But there is a consumer: you search for something, you receive something. It's an open source, open. So what exactly is wrong?
But I did say: not intelligence, but counterintelligence. That is, brainwashing the population through open sources. The opposite. But it exists, it must be accounted for. Unfortunately, there's no getting away from it. Well, about commercial activity, fine, we've all been through that. So, where the data comes from, I've already said, basically: we post it ourselves, we post it ourselves, and not only ourselves.
Theft, all the rest, that's clear, that's the classics, but again I want to say about the information that we didn't post, but that was posted for us.
Well, or just, well, even simply, look, even ordinary information security, I already mentioned here yesterday, they said, when I cracked a phone... no, I didn't; the last one was an ATM, I'll tell you about that, it's also OSINT. Look, actually, yesterday I talked to many people here, many are into this, and even when I was already leaving, I was chatting, we were walking down the street and got to talking about OSINT, and I had a case like that too, because OSINT gets its information not only from the internet. Me talking to you, maybe, information about Vienna could also be used. So, there was a forensic examination: they ordered an ATM reliability check.
Well, I arrived, it had to be done here in Moscow, and the developer of the system, the new protection system. I went out with him, we talked, took a walk, well, next day we come in, I already knew the spots where I needed to drill, and connect. In short, what did we do? I started a computer inside without tripping the alarm. Well, obviously, no need to say more, after that it's just a matter of technique. So, well, and now, after all, I've already said it: that was for the consumer, and now here's what worries me most of all, irritates me, I don't know how to put it, it's that OSINT which you can't say is open, but it exists, and has existed for a long time and constantly.
Well, let me flip through this. So, first, an agreement with Yandex. Open, not open, doesn't matter. This site collects and processes cookie files and personal data of site visitors by means of the internet service for web analytics, Yandex Metrica. By continuing to use the site, you consent to the processing of cookies under the site's policy, and so on. Yesterday, the last talk was about this too, among other things. And now cookies, well, cookies, yes, cookies, right, well, when you go online in your browser, your passwords are saved, you automatically log in somewhere, right? Now clear the cookies, what happens? Type the password again. I don't mean to say anything about what's in cookies, whether that's good or bad, just think about why every site collects cookies, is it only to bring you the information you need, after all, you enter a site belonging to, excuse me, who knows whom, including... well, here, maybe, the protection is excellent, well, further on: License agreement for the use of Yandex Browser software.
The user is notified and agrees that the Rights Holder, so, Yandex.Technologies, processes their personal data, including but not limited to... Well, read the rest yourselves. I don't mean to say anything about Yandex, that it sells information, nothing, that it does anything.
Although, it supposedly isn't obliged to hand it over to third parties. But, look, always, if we're considering information security, information protection, always remember the classics. The classics of information security. The main threat is the insider threat, it's the human being. No matter what you do, how you protect yourself, there'll always be someone who neglects it and says, "I'm so smart, I'm fine." Again, I'll tell you a case from life: so, two organizations, one has lots of money in its account, they got hacked, 64 million stolen, well, by the standards of those days. I come to another organization, a day later, also needed there, well, on a different case, we look: a secretary sits there, nails like this, a pile of tokens in the computer, in the laptop, right, a token, like this, right, and nothing gets stolen from them.
You know why? There's nothing to steal. Well, they know it from somewhere — why go on a job, so to speak, if we don't know what to steal, really. And the same thing happens here.
Will there be a vulnerability, will there be a person who passes it on. I don't want to accuse anyone, the information just appears from somewhere. Well, next, this is Microsoft. Olga Vladislavovna said yesterday that the most malicious system is Microsoft. And Microsoft itself doesn't deny it. It writes it right there in the Microsoft Privacy Statement.
There you go. Basically, if you don't agree, don't install it. Well, sorry, they say so themselves. Back in the day Bill Gates said that the ordinary American shouldn't have to think about what's on his computer. We'll decide that for him. Now they decide for the whole world.
That's a normal thing. Information is money. Accordingly, if there's information, that's money, and accordingly, it has to be used. Well, you remember the phrase: whoever owns the information owns the world.
That's how it is.
Well, and now a bit about the legal side of this. This is the personal data protection law, everyone knows it, Federal Law 152-FZ. What is personal data? Actually, a lot of ink has been spilled over what it covers, but here the definition is the usual one: any information relating directly or indirectly to an identified or identifiable natural person.
And Article 19, measures to ensure the security of personal data during processing. Well, I understand that all of this must be complied with, all this must be protected, but somehow, I don't know why, it doesn't always work.
Alas, that's how it is. Next, as I was saying, since we're touching on information security, for some reason everyone forgets the basic concepts of information security, of infrastructure security. That there is legal information protection and technical information protection. But legal information protection — well, remember: organisational measures, technical measures.
There's the organisation's security policy — again, it has to be developed. You know, once an examination came in: got a security policy at all? Well, we did an examination once. In short, a person stole personal data from a company. Well, all was tracked: how he plugged in flash drives, how he copied what, and so on. You'd think, what's OSINT got to do with it?
But the thing is, he was also, sort of, looking for someone to sell it to. And look, this information gets disclosed. And where, and how is it protected? I've said it again and again: if information isn't protected, it will, accordingly, be accessible. And once it's accessible, it can be obtained, posted, sold, and so on and so forth. But there's no escaping that. Well, I won't talk about technical information protection either, because, well, what am I going to tell people who know all this, who all studied it. But my point is different: it exists, but for some reason isn't complied with. Information security is a costly thing, yes.
But losing information is an even costlier thing.
But this is my cry from the heart: all the data gathered by Microsoft, Yandex, any sites collecting cookies — naturally, it's all protected. Because it says right here that your data must be — we accept it personally, we protect it, and we won't give it to anyone. Well yes, I believe it.
All data is protected. Where else would it go? Well, and now: OSINT isn't just gathering information, it's, really, serious analytics. Analysing that data — now that's an art. The thing is, everyone knows that 2×2=4, but when 2×2=4 applies to something, to some product, that's another matter. That is: what have you got? Are we selling or buying? What are we selling? What are we buying?
So, it doesn't matter how the information was obtained. Well, the OSINT classics: that this information has qualitative, quantitative and value characteristics. Those who've done this all know it. For attackers, basically, as I've already said, no need to collect much information, they just need to get your minimal data, at which point you — well, not you, but you know who — they can be blackmailed or something else can be done to them. And, as I said, it's all like the Field of Miracles in the Land of Fools — well, and that's the phrase I already said — but again, if they aren't aimed at some large-scale actions, again involving social engineering methods — well, remember the irreplaceable Mitnick, the book: how he got into an organisation. First I found the phone directory, then I called a department; with one it didn't work, the second said, well, my mail isn't working — yes, yes, our IT guys are doing a bad job, but you tell me what's next, send the file — and so on, it gradually unravelled; again, he got the information, he got it, the information wasn't hidden, it wasn't, so think about what ends up on the internet.
But there's something else here too: when an incident is investigated, they looked at the information, but then — how did it get there? In what way? First, I told you about the plugged-in flash drive. The person — who am I? Who had access? Who could've posted it? Well, here it's all classic information security, I'm saying it again. Well, and here again, I've already said repeatedly: is there any protection? That's the question.
Well, again, the rule: follow the rules of infosec. And countering it. Well, you've got an organisation to protect from OSINT. Well, counter-measures. Launch disinformation about your system. After all, if you recall again the classic, the classic fundamentals of information security, it's, first of all, hiding information on the informatisation object itself and how it's all protected. Well, what's there to say? Again, counterintelligence methods. Launch some disinfo, see who leaks what. Everyone's seen the Stierlitz films, and everyone else roughly knows how it works. Actually, other ways of countering and using OSINT by the classic method I simply don't see. But once again I want to stress something else.
Once again, I say: OSINT isn't what it used to be. There's the classic kind, but what's happening now — I don't know, I'm ready to debate it, this is just my personal opinion for now. But I'm saying something is happening now, because every consumer, everyone tries to find information. What's foisted on him depends on others. Well, I fit into exactly half an hour. Thank you for your attention. Any questions — I'll answer as best I can. So, the little speaker. Dmitry, I see you. —
— Good afternoon. Could you please give a definition of how exactly OSINT differs from, say, operational-search activities? I can. Look, the thing is, it's the depth of immersion. Is the answer clear? No. Using databases — I mean, OSINT exists in operational-search activities, but a different kind. I'd call it not OSINT but computer intelligence. How's computer intelligence relevant? Huh? How's computer intelligence relevant? Databases, big data. Look, OSINT for us is open sources. We kind of know. But what's in open sources? If earlier — the first OSINT I saw, back when I was still in service, we did it, it was still FidoNet. And we launched this thing and found prescriptions for narcotic-class drugs. —
— There really was correspondence there, we collected that data. It's just that in operational work, look, actually OSINT is unreliable information, it's reference information. You must never trust it 100%. But there is, let's say, in operational work, information you can trust 100%. But that's closed, special information. So that, basically, is the difference: not only open but also closed information is used. —
— Fine, but then why is everyone so fiercely trying to use OSINT? What's the point? It's quite a fashionable topic. For the last 7–8 years. Because it's been hyped, I'm saying, well, everyone uses it, everyone — I don't know, I have people here — remember, there were classes on OSINT, I gave, I just gave the data: find, collect data about me and analyse it. Well, everyone likes it, I'm saying, people here already came up, we talked, well. So we're talking about OSINT as some kind of thing that everyone likes. —
— Yes, and it's fashionable, on everyone's lips, but I'm saying, what I'm getting at, really, isn't that it exists and everyone likes it; I'm getting at the fact that information in open sources now can't be fully trusted. —
— But it never could be. I mean, I'm just trying to find out from you how exactly OSINT can be useful methodologically. Because your talk does after all somehow hint at the usefulness and use of this methodology.
Help with information security? Well, I was actually talking of something else — I talked about the information security doctrine, that people need to be prepared somehow not to fully trust the information everyone has now rushed to; and from the standpoint of classic infosec proper — well, that's all known anyway: look at the information I laid out for you; to protect yourselves, put out disinfo about your system.
Olga Vladislavovna, well, all right, of course. You know, there's also a third level of information that gets put on the internet at all. Here's what I want to say. Each of us is a specialist in some field of knowledge, knows it well. You read newspapers, interviews, whatever, on your own field, and you're amazed. Good Lord, what are they writing? Then you think: what about the rest I don't know — the approaches there are just the same, they write rubbish. So OSINT is trusting those databases, open sources, state ones or otherwise, when you need to find something, dig up dirt on a competitor, search around, or get some orienting information that you still have to dig into.
I say, the main thing is analytics.
— Colleagues, let's start with a very simple thing. OSINT is a methodology that lets you collect, analyse and verify information that is publicly available. And the basic principles we follow are that the information must be open, it must be lawfully obtained, and it must be re-verifiable. Those are the main points, and that's how, in fact, it differs from operational-search activities. That's why, in fact, people use this methodology. Not because it's cool, but because it's genuinely a method for verifying information and using it, say, in court as an evidentiary basis.
But we're not in the West, thank God.
— Look, I often act, let's say, as a specialist in court, specifically assisting in court cases. We're not at a hearing.
— Wait, we're discussing OSINT now. A real discussion's started, and honestly, that's what I was after.
— Excellent. So look, OSINT: we already know the information needs verification, but it's there. I'll throw in one more provocation: lawyers are in the room, and I'd really like them, at least at the next forum, to get up themselves and speak, because I debate with them, I've long worked with them, and I know that yesterday's problems with the questions, I understand why they came up. Same here. Thing is, sorry, a forensic examination, and generally getting info to a lawyer, to a judge, has to be in a form they understand. They're not specialists. And if you also tell them, yes, I found this information in an open source, that's it, great, so it exists. I recently brought up flight 17. For them it's yes, that's a yes. It's online, so it's correct. But is it?
Are we definitely talking about the same thing? Yes, absolutely. Literally two minutes ago I said this is a methodology that lets you verify information so that it can be re-verified in the future, including for use in court cases. Yes, of course, the lawyers thing is great. Just last night we were sitting with our respected colleague, Ms. Yulova, the younger one; we've quite often worked in this area, specifically with her. You've thrown it all into one big pile; for instance, it's unclear to me, some of the definitions you give, and they seem, I'm sorry. No, that's fine, that's fine. The thing is, that's exactly the point for discussion.
It needs to be discussed. It's a problem; once we understand it exists, it needs to be solved. What's the problem? —
— What are we discussing now? The questions or OSINT? No, let's talk about OSINT then. The problem is that you propose using the information, but it has to be verified, right? How do you verify it? —
— There are several verification methods that are used precisely for this. —
— But that's another task. Why? That's the main task of an OSINT investigation. Look, I understand, look: we found the information, and then comes the confirmation of that information, true or not. Yes, there's a lot of information, heaps of it, and our task in the process of investigation is precisely not to find information but to analyse and verify it, so that the information can be considered true. Ah, but again, look, I meant something slightly different: what if it's well-crafted disinformation? If it's well-crafted disinformation, there are always ways to verify it, to refute the hypothesis or confirm the hypothesis based on that. To plant disinfo, for example, about any of us, just go into GetContact, you know, a wonderful tool, and post from 10 different numbers that, excuse me, Yuri is a bad person.
Yes, of course, since that information is easily manipulated. But methodologically, again, the task of a specialist doing intelligence work, including OSINT, is precisely to verify the information and make it evidentiary.
— I won't even argue with that. But that's exactly where the problem lies.
— Colleagues, one small point: the roast was yesterday, and frankly, let's not... let bygones be bygones. Hallway conversations, we fully support them, but this discussion, I think, Yuri Mikhailovich, belongs on a third day of MFD, done purely for lawyers. We could set up a table. Yes, good it's there, I'm glad, thank you. And so, let's see off Yuri Mikhailovich with a round of applause and slowly move on.
12. Alexey Shulmin (Kaspersky) — “Librarian Likho — an APT group combining cyberespionage and financial motivation”
Scheduled 11:40–12:10.
Moderator's introduction
Overall, our next story will be like, well, when our next speaker pitched it to me, it will be, as I understand it, something like a real little cyber-detective story, probably, because, all in all, it's got everything. So we've got greed for profit there, and espionage as well, a proper decent little action flick. So, how APT groups operate — Alexey Shulmin will tell us that. Let's give him a round of applause. I'm sure it's going to be super hot.
Right, here you go: the clicker, the microphone.
Talk and Q&A
Check, check, check. Hi everyone, hello, hello, dear attendees, dear colleagues. My name is Lyosha, I'm from Kaspersky, a malware expert in the Advanced Threat Research Department. And today I'm here to tell you about one group whose activity we investigated. I think it'll be interesting, because it's, you know, on the one hand fairly ordinary, on the other hand it has some unique features of its own, which I'll tell you about today. So, let's go. The group is called Librarian Likho. Originally it was called Librarian Ghouls. We assigned it to the Ghouls cluster because we considered it cybercrime. But actually we later renamed it Librarian Likho, because we realized that the main motivation is, after all, cyber espionage, and the financial motivation is secondary for this group. Look, I'd like my talk to be fairly lively, so if there are any questions or remarks, please speak up, we'll discuss.
I'll add some details, throw in some more things. I hope it'll be lively and interesting enough. So, let's go. Actually, what I'm going to tell you is a small part of our big research report. It came out very recently. We worked all summer; there was a big team of authors who worked on this research. It's presented here, it's called "Notes of a Digital Auditor". It's no marketing bullshit, there's none in there, don't even hope; it's technical meat, roughly 330 pages. Here's the QR code, please grab it, it's free, read it. We looked at three threat clusters there; it's about Ukrainian groups operating against Russia first and foremost. We split them into three clusters. The first cluster is hacktivists, those who break things just for the sake of breaking them, to get some message across to their public. The second cluster is cyber espionage, the APT groups; their main goal is to get some data, find something out, hunt for some know-how and so on. And the third group is everyone else.
Lots of meat, lots of interesting technical meat, and reading this work will, overall, let you form a picture of how the adversary operates, what main techniques, tactics and procedures they use, and, generally, how to defend against it. Because the main goal of Threat Intelligence, one of the main goals of Threat Intelligence, is to know your enemy. We need to know who's attacking us, we need to know how to fight it off. So please, download the report, it's out digitally as a PDF, it's free. It's available in English too; if our customers speak English, it can be requested.
And please, have a look. Besides that, besides just the description, our report also comes with interesting diagrams, graphs built in Obsidian, which you can load, look at, click through with the mouse and move around. Tons of IOCs there, the whole infrastructure, well, the part of the infrastructure that we found, so you know what to ban, so you know which IOCs to pull into your infrastructure and block, or conversely, to search for them, in case it's already happened and you're unaware. So, now let's go, to what we're actually talking about today. It's Librarian APT, Librarian Likho APT, it's an APT, also known as Librarian Ghouls, that's how we attributed it before; our colleagues from other vendors call it either Rare Wolf or Rezet; I'll say a bit more about why. Let's take a look at what it is. It's actually an APT performing the classic role, the classic function of any APT: stealing data.
There's a sponsor behind it, no doubt, an APT, i.e. some large agency stands behind this APT's development. It's state-sponsored, so we can assume the developers have essentially unlimited resources, because, you see, however big the company countering them is, if it's a commercial vendor, the resources are still limited. When it's special tooling used by intelligence services, of course, the resources there are essentially limitless, because the stakes are very high. The main task is cyber espionage. But then again, why not steal money too, if the opportunity is there. I'll show you how these guys operate, what they do. These guys hit Russia, the Republic of Belarus, Kazakhstan. Mostly they hit industrial enterprises, to steal from them. Also, research institutes, design bureaus, so-called think tanks and universities get hit too, as we noted. Big universities at that, national ones and so on,
too. The developers follow the KISS principle, Keep It Simple, Stupid, because they don't want to overcomplicate. Actually, there may be two reasons here. First reason: we don't want to overcomplicate because we've got clumsy paws and can't write proper code. Second reason: we won't overcomplicate because, here, have some legitimate tools, now try to detect them, it'll be hard for you. Well, they think it'll be hard for us; in fact, we've seen all sorts of things, and so we've already adapted to everything; we'll detect legitimate tools if they're used in an illegitimate way. This is how they hand out their little gifts. Look, this is a real example of a spear-phishing email that they use as the initial vector. Here's what arrives; the sender may be real or may be spoofed, they can first compromise someone and then send it in their name, or simply spoof who the email really came from.
Some little attachment. Every time I speak at conferences, which happens quite often, every time I say: security awareness, security awareness, and once again security awareness. Unfortunately, it's 2025 outside, and in this room too, by the way. And people still run.pdf.exe attachments. You understand? It's funny, but they run them. They run them. It's actually a big problem, I don't know what to do about it. I'm now investigating another threat, which, due to certain internal processes of ours, I'll talk about a bit later, because first we'll release the report, then an article, then I'll talk about it at conferences. So there I observe, for example, the attacker sending a command to their implant to upload a file from the user; they want to steal a file from the user. So they grab a file from the user's desktop called "logins and passwords.xlsx".
Employee logins and passwords.xlsx. 2025, right — if you don't kill people for this, then what do you kill them for? I mean, where's the security awareness, where's at least some minimal knowledge of what's supposed to happen? And these aren't some small outfits, these are quite large, large victims. But, to be fair, I've also seen cases where victims fend off such attacks quite successfully, respond quite successfully, and, all in all, the attackers don't get very far. So this is what arrives, and unfortunately they open, open the attached archive. I just gave you an example I found, an example of such spear phishing I found, but we'll look at a slightly different gift of theirs. We'll look at this archive they hand out. Let me go to the next slide.
Ah, here. Right, I don't have a laser pointer here. Okay, look. So a RAR archive arrives as the attachment, named like this: "payment order", then 2N. That's the original spelling, I'll leave it. By the way, mistakes like these make it pretty easy to figure out who might be behind this group. Later I'll show you a few more interesting artifacts we found during the investigation — you may draw some conclusions of your own too. Inside is a file with the.scr extension. That's the screensaver extension, for screen savers, but in fact it's a perfectly ordinary MZ/PE, which, if launched via ShellExecute, runs like a regular executable. So it's a regular executable file. Well, maybe they swapped the icon, maybe not, it doesn't matter anymore. But anyway, users, unfortunately, run it. What's inside? Inside we have a classic MZ/PE built with Smart Install Maker.
Smart Install Maker is a fairly simple utility, you can unpack it yourselves, there are free unpackers, or you can knock one together in Python, it's literally a few lines, because it's incredibly simple. Inside there are three little files. data.cab is a cabinet file containing the main payload of this archive. installer.config is the installer's configuration file, which specifies what to actually do during installation of this self-extracting archive. And runtime.cab, that's some cab, I have no idea what it's for, it has nothing in it but 36 bytes of headers.
The embedded cab file data.cab contains the following files. First, it contains this PDF here — and the PDF, by the way, is genuine. It's a decoy, well, sort of a decoy, let me call it not a decoy but a red herring. That is, it's not a lure file, because by the time the user sees it, everything has already happened. It's more of a dummy file. It's a dummy file that has to be shown to the user so they calm down, that everything's fine, that it really was a PDF they opened. The PDF gets dropped over here. The legitimate executable is the curl utility. Just in case there's no curl on the system, we bring our own. And, look, a malicious LNK file. Simple as that. That's the payload. You'd think, what's there to detect here? Well, the malicious LNK file, okay, fine, that can be detected. Everything else is perfectly legitimate. An innocent PDF you can't do anything with, it's legitimate, no evil in it.
And curl — you could detect it, but that's shooting yourself in the foot. Apparently that's what the attackers are counting on. Moving on. Here's the PDF, here's the payment slip. A whole 600 rubles for life or health insurance against accidents, I think. See, against workplace accidents, we even insured someone for 600 rubles. Impressive. Well, the payment slip is real, they found it somewhere, stuck it in here as a sample, as some kind of dummy to show the user. The user sees all this and thinks, what a good boy I am, I opened the right email, I don't need it, actually I'll just send it to spam. But it's too late to flail. What happens next? Next, the install config contains commands that are executed during, as I already told you, during installation of this implant into the system. First we add a whole heap of stuff to the registry. Very noisy, very noisily a bunch of keys are added in order to install a utility called 4t Tray Minimizer.
From this developer here, 4t-niagara.com. It's a very interesting developer. Look, it positions itself as a company located in Great Britain, as a developer located over there in the United Kingdom. See, it's even been assigned an English VAT payer number, a VAT number, see, in the top right corner. A thoroughly English developer through and through. But I didn't put that image in here, because these days, you know, say the wrong thing and you'll hurt someone's feelings or discredit somebody. But you can go to this wonderful site and look at what's in the header. In the header there's a motto, well, not a motto, a slogan that instantly gives away the real developers of this real software, who think they're in British jurisdiction and everything's fine for them.
So of course this is no British developer, seriously, no, not British. Not British in the slightest, it seems to me, but these comrades keep using it. Well, okay, fine, we'll take their word for it. What happens next? Next, a command file, a batch file, is created: rezet.cmd. And that's where, by the way, the name used by our colleagues from other vendors comes from. Commands get written into it like this. Wonderful — how do we do a pause if we can't do a pause properly? We do a pause by pinging localhost, of course, as tradition dictates. And we self-delete via the batch file. When we can't delete ourselves properly, this is what we do. Right. The main stealth trick, of course, is that the directory C:\Intel, where the attackers drop their malicious toolkit, gets the system and hidden attributes slapped on, I don't know. Actually, to my mind, that gives them away, because, for instance, I have display of hidden system directories enabled by default, and when I see an Intel directory in the root of drive C that's system and hidden, I'm going to have some questions.
I'll definitely go look at what's inside it and what's being hidden from me. Well, maybe other users have it set up differently, I don't know. Yes, so much for stealth. We'll be tracking this stealth throughout all the work, all the activity of this group. You'll see, they used it a few more times. I'll point it out too. I always get a certain pleasure from finding artifacts like these and chuckling a bit under my breath when analyzing yet another threat whose authors considered themselves very clever. But unfortunately, yes, unfortunately, it is what it is, either they don't bother, most likely they just don't bother, because why would they, it works as it is, why do anything complicated. Actually, look, as an analyst who's used to complex binary implants, I miss it — give me back my, not 2007, but 2017, when there were all sorts of leaks, Snowden and so on, when there were zero-days, when there was the wonderful MS17-010, when all of that had to be researched, when there was kernel code, when it was beautiful — and now what, APTs on batch files.
And it's not the first batch-file APT. I've already talked about another APT on batch files. And the worst part, you see, is that it works. It damn well works. Users launch PDF.exe. We keep logins and passwords in plain text on the desktop. And that wasn't deception. A user really put it there like that. I don't know, just publish it then. Upload it to some cloud somewhere and post the link publicly. Come on, guys, at least something, at least somehow. No, no, no. People don't think about it. The operators use several legitimate utilities, they're listed here. Let me tell you a bit about them. First of all, there's what's called driver.exe.
driver.exe is no driver at all, of course. What did the attackers do? They took RAR, an ancient version, 3.8, I even found the executables and did a binary comparison. They stripped out all the text strings that get printed to the console, since these are console utilities. They just went and overwrote them with zeros. So the string, well, the string reference is in the code, it didn't go anywhere, but the string itself is zeroed. As a result we have a RAR that works but prints nothing to the console. It understands the same switches, works, does its thing, but the console is quiet. That's how they dissected it. The next utility is blat.exe. Yes, blat, not the swear word, it's a legitimate utility.
A legitimate utility designed for sending emails. They use it because that's how their exfiltration is set up. These folks do their exfiltration via email. They gather all sorts of stuff and send themselves a little email with findings they've harvested. Disguised as svchost we have AnyDesk, which they use actively. I'll tell you and show you how later. By the way, during the initial analysis we didn't pay much attention to AnyDesk, and it turned out to be, basically, the core of the whole attack, which they use actively later on. The same 4t Tray Minimizer is used. Well, you've got to hide the window somehow, and we can't, or we don't want to build binary implants, so we drag it along. A script is used, I'll talk about it a bit later, wol.ps1. And the Defender Control utility is used.
It's a free utility designed for manipulating Defender. In particular, these folks use it to disable that very Defender. Well, you see, these folks are all thumbs, because fixing a few keys in the registry and stopping a service is just too hard. Yes, it's too hard, you have to drag in a third-party utility and use that. Well, whatever they want. So, remember the secrecy, right? So our driver.exe, which is our RAR, is used for unpacking and packing with this wonderful password. You can see it on the slide. The password, by the way, is unique. In this respect the password is a find, because our hands aren't idle either. I took this password, combed the archives, dug well through everything we have. What do you think? The group's traces go back, the traces of this password, let's say, and it's unique, go way back to 2010. So these folks have been operating for a long time and fairly successfully.
Each of you, think about that for yourselves. I put forward the hypothesis that a password is a rather sensitive thing, and they stick with a user for a long time, for many years. A person sets a password once, then modifies it somehow, but the core stays the same, the password gets reused, and unfortunately that's a disease. So here we have one and the same password, which is a rather interesting IOC, that they've carried through the years. They dump everything they need into C:\Intel, then launch trace.lnk, which they also brought in the archive, and it runs inside 4t Tray Minimizer. Moving on. Why do the attackers need this utility? Well, it's simple, they're all thumbs, they can't properly hide an icon or hide a window, so they carry around either 4t Tray Minimizer or one more utility, which I'll get to a bit later, they carry it with them and use it to hide their windows.
Just to somehow avoid drawing extra attention from the user, though they really blow their cover everywhere. Well, apparently they're counting on those same users who launch.pdf.exe. Now look at what rezet.cmd does next. Next it installs AnyDesk on the computer in installation mode, after which it sets the password qwerty1234566 on it, so as to get access to the host without a permission prompt. So, using this password, the attacker just connects to the machine, no window is shown to the user, the user sees nothing, the attacker just connects and does what they want. Here's our Defender Control, see, it looks like this, it's pretty primitive, a whole three switches. The attackers call it with the D switch to turn Windows Defender off.
As if protection ends with Windows Defender, but no. Next, look, we call the powercfg utility 6 times to make sure the PC won't sleep, stays available, to adjust its power settings. Again, we're all thumbs, we can't do it through the registry, so we'll go this way instead. Well okay, fine. Next, again with schtasks, because again we don't touch the registry, right. A command is created, sorry, a task, a task is created called shutdown at 5 a.m. So, shut down at 5 in the morning. The name says it all, it's clear what it does. The task is needed to power off the victim every day at 5 a.m. Why 5 a.m., I'll explain a bit later. It's a rather interesting technique. It's, you know, an APT that lives by night.
What happens next? Next, finally, remember, wol.ps1 was extracted. Yes, now it's time to run it. It runs these commands, look at what's happening here. Here Edge gets launched. At least, a task is created to launch it, called WakeUpAndLaunchEdge. When I first saw this, I got tense, because, well, why would an attacker launch Edge, right? What's the point? I went to look. I thought, okay, they replaced Edge. Patched it, swapped it, something sits there instead of Edge, some little gift. No, you know, Edge turned out to be genuine. A perfectly original binary with a valid MS digital signature, everything's fine there. The real Edge. So here's a task that launches the real Edge.
What do you think, take a guess, why do the attackers need this task? Just guess. Actually it's all simple here. It's just there to wake the computer up. The computer wakes up at 1 a.m. and shuts down at 5. See how elegant it is, right? So they infected the victim, turned the PC on at night, the PC doesn't hibernate, the PC won't sleep, they worked till 5 a.m., siphoned off all they needed, the PC shut down at 5. The victim comes in the morning, boots the PC, as if all's fine, nothing happened. So, yes, and the attacker has 4 hours, enough to scoop up absolutely everything. I have no idea why they even persist on the PC, because in 4 hours you can siphon it all off. You can make proper images there, since nobody restricts you, but if the SOC isn't watching 24/7, you don't have to limit yourself on the volume exfiltrated and can siphon off absolutely everything somewhere, like to some cloud.
Let's move on. What happens next? Next our batch file deletes, in fact, curl, the curl installer, the AnyDesk installer, because they're no longer needed, trace.rar is deleted too, because it's no longer needed. We put the utility, Tray Minimizer, in the system, it's running. Next, environment variables get configured so that the blat.exe utility can do its job, so that data can be exfiltrated over email. Then the classics: registry hives are simply dumped, reg.exe is used to dump SYSTEM and SAM. I don't know, if this got past the SOC, no idea where the SOC's looking. That's not just blinders on, those blinders cover the eyes completely. You know, like covering your eyes with your ears out of shame, because to miss something like this, I don't know what you'd have to do.
What do these comrades do next? Next they figured, we already stole everything we were interested in, now let's steal some money. Look, I wrote "redacted" here, yes, I actually cut it out, but here was that same password they use, the one that was on several slides earlier. They just go through and gather into a file with the telling name wallet.rar everything they find that's related to electronic money. They'll pull out all they find, all they can reach. That's their appetite so far. And they'll grab the SAM and SYSTEM backups too, just in case.
All of this gets sent to the attackers. Think they'll stop there? No, they won't stop there. They'll also plant a miner in the system. Utilization at 110-146%, so to speak.
They download a miner, which, by the way, is legitimate, XMRig, but the pool is malicious, the controller is malicious, they just dragged in miners, now they also want the idle machine, while there's still Lyuda the secretary using the computer at 10 percent, give us the other 90, we'll put a miner on the other 90. So, I don't know why they do it this way, I don't know, seems they fear nothing, because not noticing a miner, well, again, it takes, again, blinders on the eyes, not noticing that the computer clearly doesn't belong to you anymore, that most of its resources are going to mining some Monero or something like that. So, where's the money? Where's the money? The money's gone, all of it, all stolen. And not only did they steal it, they also mined some on top.
Look, on top of that, I've seen a lot of implants, several dozen, maybe even hundreds, that I took apart while researching, while writing the report and articles, and now preparing for conferences, they use different ones, meaning they change, they don't stand still, they evolve, try all kinds of utilities, I've only talked about one specific implant, but among other things, they, you see, they also use ngrok, they may drag in the WebBrowser PassView tool, yes, they use Mipko Professional Keylogger, by the way, Mipko Professional Keylogger is quite a DLP system, but in plain terms, spyware, yes, which, among other things, records keystrokes, but the curious thing here is that MPK, Mipko Professional Keylogger, is, generally speaking, made in Russia. As I recall, it's developed in the city of Pskov, but its interface supports many languages, and in the distribution I found in the implant that I was researching, for some reason, had language settings for Ukrainian.
Well, apparently, maybe attackers somehow prefer it. That's a hypothesis, of course, but that's what these guys leave behind when they get in. So, I think that covers it all here. Yes, seems that's all. And these comrades, too. At first they only hit one computer, took all they wanted from it, and left. After that, this isn't in the presentation, I'll just tell you. They used all sorts of batch files, again, for moving over SMB, for lateral movement they moved around inside the organization, hopped to other computers, scooped it all out and, basically, left the same way. Yes, no idea why they used no binary implants in the attack. I assume the operators will watch the stream or the recording either way. Guys, make something binary already, it'd be interesting to see, I'm wasting away, my brain's atrophying reading batch files. We need something interesting, some 0-day or something original like that, because this isn't very interesting in terms of analysis, because it's all written already, you're just reading off the page.
Yes, we went looking at these wonderful guys' infrastructure, what they've got there, what domains they use. They use quite a lot of domains and servers to carry out their activities. And among other things, we found this wonderful one, look here, Yes, we're invited to sign in to a Mail.ru account. Very similar. I even compared them at the time. I opened the real Mail.ru next to it and compared them. An exact match. I blacked out a bit here, just in case. Yes, they even offer sign-in via Gosuslugi. Please pay attention to the address bar. We've got login.php there, you see? So in the attackers' minds, Mail.ru runs on standard, stock PHP. Well, guys, fine, okay, let it be so. Please pay attention to the domain name. Let me highlight it.
users-mail.ru. So they don't bother at all, and that's fine. And they think this'll fly, but judging by what they've done, and by the fact they have real scripts there that grab all of it from users, the creds the user enters, it actually works. How this is used against users, sadly, I haven't had the pleasure of observing, but at least it's there in their infrastructure. And with medium confidence we can assume it. They probably used it somehow, probably made it for something, since I don't think they worked for nothing. But apart from all that, apart from all else, the guys really don't bother much. You know, it's like back in the 2000s, when people talked about games, a joke went: "translated by the best programmers". It's the same here, you know: "site administered by the best C developers". Roughly the same thing, because the site has directory listing wide open.
You open it up, just the bare domain as it is, the bare site as it is, that's it, all this happiness is here, please, download the whole toolkit, ready to go. Well, for simplicity, apparently, so as not to bother, it's convenient to poke around the file system, take a look, we just open 127.0.0.1, and that's it, we see everything. And we've also got a wonderful thing hanging here, phpMyAdmin, right there, since they need to administer all this somehow, I don't know, either to bolt on a CMS or keep a database there. Well, anyway, we've got phpMyAdmin, which speaks to us in Russian by default, yes.
So that's the kind of evidence the guys behind this attack left about themselves. They've been operating quite a while, several years. They don't bother, they don't make binary implants, they stick to batch files. And, all in all, judging by the fact that people still run.pdf.exe, unfortunately, they do get some kind of result. I would really like the situation to change. But you understand, for our part, we'll do all we can. We detect, and we detect well. We detect actively, we detect proactively. If you read the darknet, yes, go read it. I won't advertise our product or company right now, you already know we're cool, but I'll just tell you about a few funny cases. Read the darknet, the wailing begins. Guys, someone kill Kaspersky's detection, I can't, I can't, I keep getting detected. I'll give you 3000 USDT or something along those lines to get the detection off me.
And I know the one who made that detection. And I know how well it's made, and that it can't be evaded quickly. So on our side we'll do everything we can, we're here to protect the world, we save the world, that's the mission our CEO has set, and that's how we work. But we're not everywhere. Not all are our clients. Some we won't cover, some simply don't use us, some don't use anything at all. And without a proper, adequate level of security awareness, nothing will change. They've been getting what they want, and they still are. And, unfortunately, it goes on and on and on. And at the last conference I spoke at, I think, I said that spear phishing now delivers about 80% of threats. Not at all. You know, I think spear phishing now delivers something like 90-95% of threats. And it's hard to do anything about it. They resort to all sorts of tricks, they send an archive, for example, an encrypted, password-protected one.
They put the password in the email body. But, to be fair, a decent solution, again, not plugging anyone, yes, they can pick those passwords out, plug them in, guess them, etc. They go further, they've stopped sending passwords in the email body. They email the password-protected archive so it flies past all defense layers. And the password for that archive, for example, they send some other way, some other, I don't know, some other, through some other channel. But that, too, is basically easy to detect with modern solutions. Usually in those cases the attachment is simply cut out, and in its place they put a link to the solution's web interface, where you enter that very password. After that the solution checks it, and then the solution says whether it can be handed to the user or not. We'll soon have this functionality implemented too. So on our side we're doing everything we can, but the vendor's efforts alone, which aren't everywhere, aren't enough.
First of all TI, first of all security awareness. We need to know who we're up against in order to counter them successfully. And on that positive note, I'll probably wrap up. Thank you very much. I'll be happy to chat and answer questions. Go ahead. Awesome.
Okay, I see hands, I see them. Good afternoon. Tell me, do you understand why they do it with batch files, or not? And with implants. Look, as I said, yes, I have two hypotheses. Maybe they should even be merged into one. First, they've got paws, so to speak: hard to write binary implants. Second, it's hard to detect. Hard to detect, well, supposedly, they think it's hard to detect. Actually, you know, they didn't invent this, and this technique's over a year old, over five years old, over a decade old, even. We've all seen it many times, and we're very good at detecting it. And if they think that by using a legitimate operating system mechanism, especially PowerShell, by the way, which they also love, it's their great hope. And using PowerShell, I have no idea what they're thinking. After AMSI came along, PowerShell, well, I don't know, well, fine, you might as well hand us your code directly, we'll detect it.
And the second question. Isn't Kaspersky Lab, actually, planning to make a lightweight solution to detect this stuff, with Astra, for example? Not a full-blown antivirus, but something very stripped-down that we could deploy alongside Russian solutions. Like Microsoft: here, have an antivirus. Well, look, I'm not a product manager, I'm a techie, so I really can't commit to which products will be made. Yes, but there is, for example, KFA, you know, yes, Kaspersky Free Antivirus, great, please use it. There are trial versions of our products, please install them. If something's already happened, of course, you need to get protected.
We provide the full range of both services and products for protecting the information world, the digital world. We can cover basically anything. We can help with everything. Please, come to us, we'll do it. And on top of that, we really do care about what we do, we put our heart into it. We're all into it, there are no random people here. We're all professionals who are engaged, who are ambitious, who are interested in growing their expertise. And so all the new stuff the attackers roll out, we see it, often we see it even before the attackers can start using it. That happens too, and then they're very surprised. How come? They haven't even used it yet, haven't even delivered it to victims, and Kaspersky somehow already detected it. How does that happen? I don't know how. Somehow.
Somehow. They're generally pretty naive people, they use services like VirusTotal, well, not VirusTotal, of course, but similar ones, to check that their little creation isn't detected by anyone or anything. Now I'll finish it and send it out to users. And when they see it's supposedly undetected by anyone or anything, they send it out to users hoping that's really the case, that we have no other engines. We won't disappoint them. We do, of course, have our own technologies. You see, it's a cat-and-mouse game that will never end. We catch them, they run, we catch them, they run again. All of this, of course, has to end not with technical, not only with technical means of countering them.
I always talk about this, and I'll probably tell it again now. When we watch the evolution of a fellow who has decided to try a black hat on, try it on, put it on, not for himself, on himself, whatever. Here's what we observe. For a while at first he's terrified. He fears they'll come for him, that tomorrow at 6 a.m. they break his door in, put him face down on the floor, and it all begins. And then time passes, weeks, months, maybe years, and nothing happens, and nobody breaks down his door, and nobody puts him face down. And he figures everything's fine, that's it, he can keep blackhatting, keep working. Some, the especially reckless ones, start working against Russia while in Russia. It all ends exactly the way they feared. Work the RU, and they come for you at dawn.
That's how it goes. But, sadly, these fellows don't stop doing it. Well, as they wish. They all hope for Art. 272, 273, 274, but it's not always so. Sometimes it turns out a little differently and get very, very long sentences. —
— More questions, please. Alexey, hello. Thanks for the talk, very interesting. But, unfortunately, nothing was said about the victim computer. What system was on it, was there any protection, and why didn't it work. Because usually, if it's Windows and Defender is there, it should always trigger on a BAT file and on a PowerShell script. Well, if it's up to date, of course. Thank you. Thanks for the question. Actually, it needs splitting into sub-questions. I'll do just that and answer them one by one. First, we don't see everything. Let's start there. It's not because we're blind, but because part of the data is simply cut out. You know, we're bound by regulators' requirements, bound by all the GDPR stuff and so on. So quite often, as a rule, we don't know who the victim is, who's there. We just see anonymized information and nothing more. We have nothing that would let us attribute it to a specific victim.
We just see what's happening, we see the facts, but don't know who it is. And that's how it should be. And for that we've built a special system, a mechanism, so that all this data simply never reaches us. It's all in our KSN agreement, you can read it. Secondly, who told you, where did you get the idea that it wasn't detected? Maybe it actually was detected. Well, I mean, look, most often, when we start studying some story, some statistics, it all started with some detection, with some activity that looked strange to our products. You know, something's going on, and our product goes into suspicious-dog mode, squints like this: something's off, need to take a look. These are fairly complex technologies, I probably can't even recall them all from memory now, but it goes as far as, depending on the conditions the binary runs in, for example, our emulation level changes, we can emulate deeply, or we can emulate not very deeply, because deep emulation costs resources, on the one hand, but on the other hand, deep emulation lets you pull out what's hidden under five layers of crypter, in other cases we don't emulate very deeply.
But if we're seeing a fairly detailed picture, as a rule, it means something happened, there was a detection. Of course, we protect against all this, no doubt about it, don't even think about that, we simply have strict protocols. Look, before I can release a report or publish an article, or talk at a conference, I must check that every single one, absolutely all the implants that were found, are all detected, and detected well. But that's also a catch-up strategy, which in itself is a bit flawed, because when you're catching up, you'll always be catching up, you'll never get ahead, and we're trying to get rid of that flawed strategy and not use it, not run our strategy reactively. That is, our job is to cover users, to protect them from threats that haven't come yet, haven't happened yet. I'll give you a simple example. One of the common threats is, for example, ransomware.
You all know it. When they encrypt data, then extort a ransom to decrypt it. Our products implement a proactive detection system, where the product just watches what's happening in the system. If it sees that one file in the system was changed, its name changed, a second file changed, its name changed, entropy went up, say, a third, fourth, fifth, after that the product says: OK, that's enough. I don't know this threat, but I think it's a threat. It'll kill the threat, roll back all the other files, the ones that got encrypted, restore them, and for the user it'll be completely transparent, if, for example, it's KES. The user won't notice a thing, the ransomware is just shot down on takeoff, everything it damaged will be restored, and the user just keeps working as before. So we protect, and we even try not to bother the user unnecessarily.
Often the user doesn't even know what happened to them, and we already protected them, covered them, all fine. So we protect, we protect. As for Defender, you saw, they turn it off in this attack. If the attackers have enough privileges, they'll just switch it off, and that's it. That's one side. On the other, well, what are they doing? They install stuff via batch file. Install via batch file. Then start collecting. Well, they copy files. You'll agree, a file copy request can't be detected. That's legitimate activity.
More questions, please. Hello. I wanted to ask, couldn't this thing, their stupidity with the password, be pushed further? Like, send a request to the social networks, say, VKontakte, to their archive, to check which user had that password. Or Mail.ru. Look, did I understand you correctly? You're suggesting doing OSINT on what? No, no, making an official request to Mail.ru or VKontakte, so they check which user once had that password. Well, look, I have no idea whether that's possible or not, because I don't talk to the colleagues who provide such social services, but it's not our job, we don't do that, I mean, we don't investigate incidents of that kind. That is, we have incident response, but we're not a government agency, we're a commercial company. Our job is to protect, we protect users.
On the one hand, it's our mission, that's how we work. On the other hand, sometimes it's a bit, you know, oh, right now I'd..., right now I'd... — no, you can't, we protect. We don't attack, we protect. And even when you see some vulnerabilities in the attacker's infrastructure or something else that could be used to your advantage, we don't do it, because we're about Defensive, we're not about Offensive, we're about Defensive. Got it, thank you. —
— Colleagues, I know many questions remain, but Alexey is staying with us at the venue, and we need to move on. Let's give Alexey a round of applause. Thank you very much.
That was great.
13. Vladislav Azersky (F6) — “UWP in the DFIR crosshairs: what modern Windows apps are hiding”
Scheduled 12:15–12:35.
Moderator's introduction
Most often, what ends up in the crosshairs of DFIR specialists is all kinds of different things. But our next speaker has put ordinary Windows applications there. How it all works will be explained to us by Vladislav Azersky from F6. Vladislav, please, let's support him with a round of applause.
Clicker, microphone.
[applause]
Talk and Q&A
Hi everyone. Well, I think we can actually get started and talk, for the most part, not even about incident response so much, but more about computer forensics.
Today we'll be talking about, in fact, applications that we often come across in, say, the Microsoft Store, but not just there. We can even build them ourselves, and they'll be packaged the same way as a UWP app. That's exactly what we'll talk about today. A bit about myself. I've been working in information security for 7 years now, mostly doing information security incident response, digital forensics, and I very often run various cyber exercises in the Purple Teaming format. I also like taking part in various conferences, and, accordingly, doing research on both Offensive Security and Defensive topics.
So what is UWP? We can actually start at the very beginning of the story, when it all began. Windows 8 introduced various applications that were called Metro Apps, or Modern Apps, and they were, they ran, so to speak, as containers, isolated, and they had no access to any important system resources of the OS itself. And that's the first part. Then, some time later, Windows 10 came out, and that's where Universal Apps appeared, along with the UWP ecosystem itself, which let us create, essentially, universal applications that could be run, say, on some IoT device where we have a Windows operating system, as well as on a regular Surface Hub or a regular PC.
And at the same time we wouldn't need to recompile all of it, rebuild it, which is nice. But keep the following in mind. Supposedly all of this is isolated. We mostly see these apps, for example, in the Microsoft Store. But let's look at what else Microsoft did. In a certain build, 1607, they implemented a feature called Desktop Bridge, which let us package ordinary Win32 programs as UWP. And at the same time we could set certain permissions that allowed our app to freely access the system components of the OS itself, the functions of various APIs. That is, in this case the isolation was gone. And that's a very important point.
So, how are they represented at the file level? Most often, if we're using the Microsoft Store, we can simply download an app by clicking a button, and we won't see any file that gets installed, we'll just see a progress bar. So what are they, what do they look like? Today, this is an MSIX package being installed. They're essentially a kind of archive, they let us simplify the installation and update process for various apps, in this case, like a regular installer. It's an evolution of APPX, another format that was tied purely to UWP, which had that whole isolation business, but that was dropped later.
What else is important to know? Here on the screenshot you can see, I've highlighted three main files that we'll find in an MSIX package. First is AppxManifest.xml, which spells out how the application will be installed, how it will be launched, and also what permissions can be granted to it. For example, run full trust, which means it has access to all the components of the system itself, like a personal Windows system. Next is the signature itself, because every MSIX package must be signed with a valid certificate. We'll talk about that a bit later. And in that file, in the package, the certificate is also present.
And besides that, another very important thing. It's the BlockMap file, listing all the files in the package with their hashes, which can actually help us.
Let's take a look at how they're classified. We basically have four types of applications, three of which we can find in the Microsoft Store, the last one is signed by the developer themselves. First, system applications, essentially part of the Windows build by default, you can't, say, update them with a button, I mean remove them. And they get updated, for example, through the Microsoft Store. Next we have applications that are developed by Microsoft itself. That's, for example, something like Paint, or Calculator, which we all see often. And these are basically the apps that also often ship with the Windows build, but we can remove them, reinstall them, do whatever we want with these apps.
The next two items are, accordingly, applications that are in the Microsoft Store but aren't owned or developed by Microsoft. And here these are apps like, for example, Netflix, which we can easily install through the Microsoft Store, or the well-known Python interpreter, in various versions, which we can grab from the Microsoft Store and install without any problems. And the last, most important one, which we will encounter, and actually have sometimes seen in various incidents, is an application signed by the developer. And here I made a specific note that most malicious MSIX applications are signed precisely with a developer certificate. This may be a case where the attackers compromised some company, collected their certificates, and then signed their own MSIX package with them.
Why is this needed at all? If we try to install some MSIX package and it isn't signed with a valid signature, what happens is that when the user clicks the Install button, they'll see: oh, sorry, we won't let you run this application, because we couldn't validate your signature. And that's what we're going to talk about. How are MSIX files created in the first place? How can we create them? What software can we use? Here there are actually two main giants. One is the MSIX Packaging Tool, which can also be installed through the Microsoft Store. And also a commercial product, Advanced Installer, which also has a free version that lets us take our regular application and package it as UWP.
Okay. But actually, can we create one without even using this tooling? Yes, in this case we can take the MakeAppx utility. And for this we actually need two main things, maybe even three. First. We have to pick the application we want to package. Second. We have to write the AppxManifest.xml file by hand, which specifies how our application will be installed, how it's launched and what permissions we'll give it. For example, we want to let this application run without isolation and have access to all components. And for that, in the second note, we can specify runFullTrust, which lets the application avoid being isolated in some container.
After we've created our package with MakeAppx, we actually have to sign it. We can go the following route: create a self-signed certificate and sign the package with the SignTool utility. But keep one thing in mind here. Since this certificate is self-signed, it won't actually pass the checks against the various trusted certificate authorities, and before installing the MSIX package itself, we'll need to import the certificate into the system, which we created, into Trusted Root under Local Computer, which will actually require admin rights from the user. But, as I said, we supposedly can't install unsigned MSIX packages, yet with the arrival of Windows 11 the developers decided to do the following: we can easily, using a certain PowerShell cmdlet, install exactly these unsigned applications. To do that, we'll need to specify in the manifest, in the Publisher field, the Organization ID that's shown right here.
It will always be the same, and so this will be one of the indicators to find unsigned applications later on. What can we look at to understand what's going on out in the world? For example, in mid-2023, Microsoft, or rather their Threat Intelligence division, recorded certain activity by some financially motivated groups. And they were distributing MSIX packages, and to launch their installation they used the ms-appinstaller handler. It ended up being disabled by Microsoft some time later, because they saw that, yes, there was a lot of both phishing attacks and distribution via SEO poisoning, where we, in this case the attackers, push their websites up to higher positions in the index of, say, Yandex, Google and so on, search engines. They also distributed it through various ad placements. And, for example, they also used phishing via Microsoft Teams.
How did the whole story unfold here? There's a certain CVE, and honestly, I never found a more detailed description of how it works, but I can assume that in this case the attackers signed with, say, invalid certificates or signatures, but thanks to this vulnerability, they could pass through without issue, and the user got this window where they could hit the Install button. But at the same time, the publisher shown was not at all the company that, so to speak, developed Perimeter 81.
What else to remember? And yes, as I described, there are some financially motivated groups that have looked at this whole thing and used it, but there aren't that many cases. What else do we need to keep in mind? There is a framework that, for example, the FIN7 group used, and they, what did they do? Created an MSIX package and embedded this framework, which lets us, without touching the program's actual source code, for example, run scripts before it executes. And we see that for this we need to write into config.json our PS1 script, for example, or a batch file, really any script at all, and then it will be executed before what we actually need runs, for instance, the application.
So what can we actually see from the DFIR perspective? What we'll mostly care about is the story of which applications were actually installed on our system via MSIX packages. And here there are actually four main points. The first one is the directories, because when an application is installed, a specific directory with its name gets created, plus three other sources of information that we'll talk about now. In the Windows operating system there is an SQLite database called StateRepository-Deployment, and in some of its tables we can find quite useful information. If, for example, the application is installed the way it was in those attacks carried out by financially motivated groups, we'll be able to see the link, in this case the URL, from which our MSIX package was essentially downloaded. And next, if it was simply launched with a double-click from the operating system, we'll just see the full path to that file.
But beyond that, if you dig deeper into the various tables, we can even see the hashes of every file that was in the MSIX packages. And consequently we can also, for example, having some indicators of compromise, search for certain files, or simply collect those indicators of compromise. But you need to keep in mind that this database and the next one mostly contain only information about installed applications, but not information about removed ones. But actually we don't really need the removed ones in this case. The next database is also SQLite 3. It also has a few useful tables, which are listed here. And what do we learn here? For example, I said that Windows 11 lets us install unsigned MSIX packages. Consequently, within one table we can find this Organization ID, which will tell us that yes, the package was installed without a signature.
And in addition to that, find out who the publisher is, which is actually also important. Besides that, we have everyone's favourite event log, which exists in all Windows operating systems. And here I broke down four main cases. The first and second cases are based on the hypothesis that the attackers installed the application using the PowerShell interpreter. In one case a signed package was installed, in the other an unsigned one. And the difference here will mainly be that when an unsigned package is created, event 9545 gets generated, which explicitly tells us that the package is unsigned.
Next, the popular options are when the user simply double-clicks the icon of the package, and it launches. Or using ms-appinstaller. As with the typical cases, the difference will be that in one case the user simply launches it, and we'll see the full path to the file, while in the other it will point to the source link from which the file was downloaded.
Besides that, we also have logs such as the App Installer log. What's important here and what to keep in mind? I said earlier that we can look, for example, in the event logs and see that with PowerShell some signed or unsigned packages were installed. Well, App Installer lacks that information, but it has the third and fourth cases, when we launch it either through the graphical interface or via the ms-appinstaller handler. To tell them apart, we can simply look at the line that is highlighted in green here, and in this case understand how the launch was performed. But besides that, there's one more thing. What if the attacker needs some applications from the Microsoft Store? In that case they'll most likely use the winget utility, which will let them install some package. This doesn't even require administrator rights, and it can be seen right there in the log.
Okay, we've talked about how attackers can deliver some MSIX packages, use some CVEs to launch them without any trouble, import a certificate and all the rest. But sometimes this even requires admin rights. So let's think: what is there in the Microsoft Store that can be installed without admin rights and that would be of interest to attackers? Here I singled out five specific areas. The first is various interpreters. In this case, in the Microsoft Store it's the Python interpreter and the Julia one. Then there are tunnels, which let us build, roughly speaking, a tunnel between our host and the infrastructure, so as to, roughly speaking, have access to the victim. Third is a web browser. For example, some attackers like, as part of an attack, to deliver their own portable versions of browsers and then use them to browse web resources located inside the company.
Fourth is SSH/RDP agents. And the last is the ever-popular Remote Access Tool, or, in other words, legitimate remote administration utilities, which are exactly what lets us help some accountants and so on. In that case support helps them, connects via TeamViewer and then works through and fixes some technical problem. Okay, let's talk about interpreters. I'm looking at interpreters here and, in this case, tunnels. What can we do with interpreters? Let's say the attacker has managed to somehow deliver some payload, or they already have access to the host; they can use the winget command.
The command line is shown here. After the interpreter is installed, they can, for example, either pass the Python interpreter, or Julia, a file with code, or simply specify some payload on the command line. This payload essentially gives the attacker a CMD for interacting with the host, and that's our reverse shell. And the second very important thing is, basically, tunnels. With tunnels the situation is interesting. In almost all cases involving various financially motivated groups that drop ransomware into the infrastructure and run it, they use tunnels.
The most popular, which I also found in Microsoft Store, recently added, are actually ngrok and localtunnel. They can also be installed with winget and then launched. But very often attackers also establish persistence for these utilities. And the very last one, which I saved for the end, is VS Code. VS Code is generally a story about, actually it's a developer IDE, i.e. their environment. And when we install it, we actually also get an additional utility, code.exe, with which we can create a tunnel. But for that we need to authenticate with GitHub. Most often that's no longer a problem for attackers.
And we move on to the main, final block: recommendations, at least the minimum that can be done. First, prohibit users from installing untrusted packages. In this case, packages that were not delivered through the Microsoft Store. And second, prohibit users, unprivileged ones in general, from launching or installing any MSIX packages, even from that same Microsoft Store. So they'll need admin rights for this. That is, at the very least call an administrator to do it. What else can be done? Write detection rules, for example, for a specific elevated cmdlet with a parameter that tells us that an unsigned MSIX package was installed.
And also, for example, collect event 9545, in which, as we also saw, there is an indication: oh, we've got a package installed here. And lastly, if we're saying the company has some SIEM or log management systems, we can simply collect, for example, these four events listed here, and additionally, on top of that, collect various commands, in this case the PowerShell log and the event log, which is also very important. That's actually all. Later I'll even publish the presentation with additional material, which will additionally cover some of the questions from this presentation.
Colleagues, your questions, if there are any. As always, you've broken my whole audience. It happens, it happens. Right, if there are no hands, Vladislav, again, thank you very much. That was super cool.
So, dear guests, we're now going on a short break and we'll meet here again at 1:10 for the second session of talks. Thank you.
[A break (12:40–13:10 in the program) is cut from the recording; timecodes run without a gap.]
14. Nikita Vyugin (MKO Systems) — “The value of security tools: what we wouldn't have if we had everything”
Scheduled 13:10–13:40.
Moderator's introduction
About the golden antelope. You'd think: strike a hoof, and gold will pour out. But each time it gets heavier and heavier. And as if it's no longer a blessing to us. It's roughly the same story with security tools. You'd think, if we had everything, we'd be invincible. But is that really so? My much-loved and much-respected colleague Nikita Vyugin will tell us about that. Let's give him a round of applause. Nikita, the floor is yours. —
Talk and Q&A
Hello everyone, colleagues. As Dima has already confidently said, well, just in case, yes, I'm Nikita, I work at MKO Systems. I think we already know each other. If not, I'm glad to meet you.
It's great to see you all here, but let's move on to a rather interesting story. As Dima correctly framed my topic, today I'll be talking about the value of security tools.
From both sides. I think a bit later you'll understand what I mean. I'll start, actually, from afar. Preparing for today's event, I wrote 33 pages of text, terribly complex, dripping with technical detail, descriptions, abbreviations, with modern security tools and all sorts of other things. Overall, I wanted to present it to you. And then I was lying on the couch scrolling through Telegram. And I saw this text in Dmitry Boroshchuk's channel. Dima, sorry if anything, do subscribe, it gets extremely interesting there. And this text really struck a nerve with me. It struck a nerve because, indirectly, I belong to the cohort of marketers, specifically in information security.
And somehow it all made me really angry. Like, Dima, we're working here, and you go and write things like that. We're all for maturity here, real maturity, I mean. Pity nobody held up that little sign for me, because apparently I was tired by then.
I still want to discuss security tools with you, with their nuances. But above all, I want to discuss them from two positions I think matter right now. Especially important for companies that are only starting to build some more or less serious security function in their organizations. The two positions are,
However strange it may sound, the position of reasonableness and of maturity. Too early. Actually, if we're talking from the position of reasonableness, then in my subjective understanding, reasonableness isn't the moment when we look at something and go, damn, that looks reasonable. Reasonableness is when information security, as a department of some company, chooses a security solution for itself, consulting the other departments. Not just with the business, as usual, not only, well, we've passed that stage, that we don't build fences the business can't get through. Here we also bring in physical security, the security service, the economic security service, lawyers, HR, and countless others. That is, modern security systems are deployed into processes that affect every department. And deploying something without knowing how those business departments work, we make a big mistake. And sometimes we just waste money, as often happens.
In terms of maturity, it's even simpler. Dima's absolutely right in what he said: maturity isn't the ability and willingness to buy any tools for millions upon millions. Maturity is clearly understanding how we build them, whether we need them or not, and how we integrate them in-house. Once again, my task today is to remind you of the familiar security practice tools, and without all this marketing tinsel, in order to step back a little and remember what these things are actually for in the first place, at what stages, to think about at what stages we might need them, or rather you might need them, and whether to deploy them, and what it costs us, not in money terms, although in money terms too.
Let's start with scary antivirus, as ever, our familiar one, our most, most, most basic level. Actually, for those who aren't in the know, let me remind you that an antivirus, in essence, simply removes malware… blocks and removes malware on endpoints. It has a catalog of all sorts of critters and a catalog of suspicious operations. If it finds a match on one or the other, it, essentially, stops the process or kills the file.
What's important to grasp? Antivirus is the most, most, most basic level of endpoint protection. This is all about users, about office machines, about servers and so on. It's the very first step in cybersecurity. What's important to know even when deploying an antivirus? Every information security tool has its nuances. And speaking of nuances, what will we run into when deploying an antivirus? As always, false positives. Triggering on things it shouldn't trigger on. On some legitimate processes, on some legitimate files. There'll be some. You need to be ready for that. A false sense of security. As often happens. Buy our antivirus and everything will be fine. That's not quite so.
Up to some level, everything will be fine. And in its basic form, again, I'm talking about the basic level, it integrates with other security tools rather peculiarly, and most often super basic products simply don't integrate with anything at all. Another important point is the BYOD policy in companies. That's bring your own device, when our entire infrastructure is covered by antivirus, and then a remote colleague comes along with some personal device of their own, knocks on our infrastructure's door with no corporate antivirus, and, accordingly, all sorts of wonders can start happening to us. And, accordingly, one more point. Everything the antivirus takes out, it takes out into quarantine or for good. And sometimes it's quite hard to investigate incidents after that.
But you have to understand that despite some seeming difficulties and limitations, I subjectively believe that if your company has at least one digital device in today's world, you need an antivirus. Since everyone in a company has at least one digital device, you need an antivirus.
— Without an antivirus, the risks we'll face are that 90% of simple threats won't be blocked right away. That is, all these critters living online, which get onto USB sticks and so on, will be wreaking all kinds of havoc in our infrastructure. Without an antivirus, if we already have some security structure, and suddenly a strange thought strikes you: maybe I should drop the antivirus. No, don't drop it, because it's that first barrier stage. If we remove it, the load on the rest of the infrastructure goes up. And, accordingly, that other infrastructure usually costs more. So, if the load on it goes up, we'll, accordingly, be paying more for it.
Well, accordingly, there are some basic levels of compliance with various requirements. And as consequences, for the corporate story we always have financial risks and reputational consequences. I'll say it again: antivirus in today's Russia is perhaps the integral, most important and most initial and non-negligible part of information security. Next, let's go in ascending order.
DLP. I don't think it needs much explaining either. In essence, it simply controls potential leak channels. It identifies sensitive data by signatures, by content analysis, by document classification, by security labels and so on. And, essentially, it prohibits copying, encrypting, moving these files somewhere. Essentially, DLP is used, to the point that, whereas an antivirus, I believe, should always be used, DLP is used in cases where we have a high risk of leakage of some insider information, trade secrets, or we need to comply with some regulation.
Let's imagine a simple situation. We have only 15 machines, but we deal with some kind of private data, some personal data. And it would seem, with 15 machines in a small little firm, you could get by on antivirus alone. But let's add to that, say, some tyrant of a boss who really likes to periodically part ways with our employees on bad terms. Well, that's just how he is. Do we need DLP in that case? Well, I think you know the answer. So, it would seem the scale is small, yet DLP is needed. What's another small point about DLP? DLP is extremely sensitive to configuration.
This is important. Extremely sensitive to data classification. If we configure DLP badly, we'll get a lot of noise in the metrics that we'll have to respond to, and that will make our lives harder. Because, essentially, if we overtighten the settings and crank every knob to the max, then we'll simply be drowning in that noise. We'll also need to integrate it with our business processes, because just adding DLP as an afterthought and saying, go work, no, that won't do. When you scale, it will need further tuning. When the business processes scale, it will need further tuning. So it's not a plug-and-play story where we plug it in, it runs, and that's it, it's wonderful.
Plus, of course, there's the balance, which is also written up here. We need to set it up in a balanced way so that, again, it doesn't hinder business processes too much, doesn't kill our business processes, but lets nothing extra through either. This thing isn't constant; we'll have to keep adjusting that level. And for that we'll need proper specialists. About DLP. What awaits us without DLP? A growing risk of data leaks.
I've seen different statistics. The figure I believe most is that up to 30% of data leaks are cut by using DLP in a company. Some say 90–100 in the marketing brochures, but as they say, necessity's the mother of invention — they'll get it out somehow. But that core 30% we cut off just fine. Accordingly, we won't be able to control leak channels, we won't be able to properly investigate incidents, if we don't know where and how it all leaked, where it was all stored, and who had it. And, accordingly, if you're part of some kind of supply chain, then complying with the regulations agreed between the two companies will be extremely difficult.
Let's move on. And next we'll be talking about EDR and XDR. These are quite interesting things, but I've combined them on one slide. And with EDR everything is fairly simple. Well, relatively simple. EDR is a kind of overseer. It monitors endpoints, PCs and servers. It does it a bit more interestingly than an antivirus, in the sense that it records process behaviour, network connections, script execution. It's not just matching an antivirus database and some database of actions.
What do we need EDR for? We need EDR to detect sophisticated threats and to have the means to investigate them afterwards. I think you all know it, but just in case I'll say that EDR is implemented as follows. It's an agent that's deployed to a machine and then, in fact, runs there. Cool thing. It's often used, and it makes sense to use it if you understand that you're in the risk zone for more or less sophisticated attacks.
XDR is essentially an extended EDR that uses as its input data the network, email, and clouds, and provides a kind of correlation of events with each other. XDR combines data, applies analytics, often applies machine learning and correlations. Also a very, very cool tool. The difference between them. EDR works on endpoints. XDR covers the entire infrastructure, including networks and clouds.
Nuances of working with these two tools and deploying them. As always, noise, alert overload. There's no getting away from that. I think it's inherent in practically every security tool that has any ability to alert. Complex deployment and configuration. That's really so.
An important point that I think is worth highlighting. By deploying EDR and XDR, you've reached a certain scale. And that scale will grow. Accordingly, you'll need new licences, new endpoints. It's important to understand that if you go down this road, then you'll also be scaling your spending on these systems. These systems are cool, but expensive. So if you're told that you'll install EDR and all your problems will vanish, that's probably true, but you need to look, count, calculate, and make a decision for yourself. How important is this to you — if you need it, deploy it. If you "need" it because some slick salesperson came in and sold you the idea, well, do the maths. Count, count, don't be lazy. It's important.
And, as an afterthought, what awaits us without EDR, by contrast? So with EDR we understand: we defend against sophisticated threats, we get a certain headache in money, a certain headache in scaling and a certain maintenance complexity. But nevertheless, without EDR and XDR we have poor visibility of endpoints.
We detect incidents late. That is, meaning that the earlier we detect an incident, the easier it is to work with, the easier it is to investigate, the easier it is to prevent, and the more we reduce the consequences of that incident.
So without EDR, growing damage from attacks awaits us, for the reason above. And, of course, non-compliance with many modern security and protection standards.
I'll repeat my thesis. Powerful tools, but you need to prepare for deploying such tools. This is no longer a basic antivirus, no longer a basic DLP system. Here — though a DLP system without proper configuration is quite a headache too, but here you'll really have to think hard about everything and do serious estimates. And be ready for the fact that you have major changes coming in security at the very least, and in IT for sure.
Next, SIEM. My favourite thing, about which there are a lot of misconceptions, at least from my experience of talking to people at various events. SIEM, in essence, can be thought of, very simply, as an alarm system. When something goes wrong in our infrastructure, SIEM screams and, in fact, collects some logs, stores events, integrates with a huge pile of different devices and different information security tools. In skilled hands, SIEM is an indispensable assistant, because at scale, keeping track of everything, if you have a fairly large infrastructure, manually keeping track of everything that happens in your infrastructure, well, that's impossible, frankly, impossible.
We'll either miss everything, or we need a battalion of security people running around by hand. Well, that's just ridiculous.
But there's one important nuance. For your SIEM to work adequately, you need to have aced all the previous stages and some of the following ones. Otherwise we get some very unpleasant things. I'll explain which ones later. And one more important point. SIEM doesn't respond to incidents automatically in any way. It doesn't protect on its own. It detects and informs the operator about what's happening in the infrastructure. It's not SOAR, not some kind of artificial intelligence. It's simply this: SIEM screams, and the IT guy cries. That's the story.
On the nuances, first, what's important to say — if you remember, I said that everything we had before needs configuring to perfection. Because with SIEM the main principle is this. Garbage in, garbage out. If all our systems are spewing some nonsense, some noise, our SIEM is just as noisy and, in fact, nothing good comes of it. Second point. For every infrastructure, for every individual company, sometimes per business unit, you'll have to write detailed SIEM behaviour rules, when it will scream. There are no ultimate ones, there are some, let's say, presets, but they all need to be fine-tuned, they all need thinking through with the business, since what's open to accountants isn't always open to IT people and so on — so there are important points.
And I probably shouldn't forget to say something about reasonableness. Well, there's the classic example. 6,000 machines, 6 IT specialists.
Do you have… You roll out SIEM. Are your 6 IT people ready to maintain 6,000 machines under a poorly tuned SIEM? Well, I think the answer is obvious too.
The nuances, the difficulties of deploying SIEM. What are SIEM's advantages? Or, more precisely: what awaits us without SIEM? No unified security picture. We collect everything on separate monitors. That is, we don't aggregate data across our systems anywhere. Adequately, I mean. We don't see whether an attack is unfolding in real time, or we see it through the fog of war, which is quite difficult. The SOC, if you have one, becomes relatively blind. As for DFIR, probably no point even mentioning it here. And, as always, non-compliance with standards, if you're in some field where a certain level of standards must be observed, or if you're in a supply chain with some company that imposes certain requirements. And these days that's not rare, as far as I know.
In general, what can I say about SIEM? And in conclusion: without SIEM, a company is left without central security monitoring. All the other systems can work, and can even work well, but each on its own. That is, we have no way to correlate and aggregate data in one place. We won't have a complete security picture. And if we don't have a complete security picture, some kind of unified one, let's say not a complete, but a unified security picture, then even if we have a huge number of great infosec specialists, we're going to lose a lot of time and resources trying to somehow respond to incidents and keep them under control.
Next up for us is SOAR. Probably the system I personally find most interesting of all.
SOAR, the way I see it, is a kind of automatic problem-cruncher. We integrate it with the SIEM, with the EDR, with all the other systems, write playbooks into it — read: simply scripts of automated response to specific triggers. And when those triggers fire, it goes and runs the playbook.
Now SOAR is probably the one tool that we absolutely must integrate through negotiations and agreements between all the departments involved in the business, and not just the business. Otherwise it seems to make no sense at all. SOAR, unlike a SIEM, doesn't collect data on its own, but it automates the response. Again, you need to understand that SOAR is not artificial intelligence. It doesn't invent anything new. Whatever playbooks you write, that's what it'll follow. And accordingly, it won't solve all your problems. But with the right approach, it will do a decent job of the part of the work that otherwise has to be done by hand, and solve problems at the outset. And accordingly, your own infosec people will find it easier to work from there.
The caveats. Obviously, with systems like these. Integration is complex to set up. It will be very hard, fairly hard, to develop the playbooks, the automated responses. It really is... And again, don't forget, we're scaling all the time, our threats keep changing. These playbooks are one thing today, tomorrow they're modified, the day after it's a third version, a tenth, a twenty-fifth. There's the danger of automated actions. Meaning playbooks must be written with your head on. Sometimes automated actions are written in such a way that they'd better never have been written. It runs into something, some trigger, it fires, someone didn't think it through or write it properly, business goes down.
That happens too. Again, what we run into most often, an important point — resistance from other teams. Other business units don't always trust infosec, don't always trust some new mechanisms, don't always trust some new processes, and can throw a spanner in the works. This is about how, for some chief accountant with a USB stick lying in the top desk drawer with the keys sitting on the desk, it'll be hard to explain we'll now install DLP on you, which will monitor your every action, and SOAR on top of that, which will automatically do something or other.
Not everyone's ready for that, so you end up having to do educational work. Well, and as I said, SOAR doesn't think, it does what it's told. And if we've messed up, we can't pin it on SOAR. Here you only have yourself to blame. But on the flip side, as always, without SOAR corporate infosec is left without automation and standardization of response, which matters at scale. When, again, we have 15 machines, a decent antivirus rolled out, DLP configured properly, and overall we've got an IT guy and an infosec guy, if we even have them at all, who can run around on foot, it's all clear, all good. We don't especially need it, really. And automation and standardization of response are, in a sense, already covered. But when we have 6,000 machines, we'll have to respond somehow.
SOAR doesn't replace a SIEM, or EDR, or XDR. It speeds things up and makes life easier, including the lives of SOC analysts. That's its main story: to offload the SOC guys, turning them from a fire brigade that's always running somewhere, tripping and dropping things along the way, into an effective response center.
And actually, since we're on the subject of SOC, let's talk about SOC.
SOC. You can't really call it an infosec tool in the usual sense. A SOC is a sort of center for monitoring, detection, and response. It's a team of specialists plus infrastructure plus tools. Essentially they gather signals from these systems, triage them, analyze events, escalate incidents, kick off some kind of response, use playbooks, TI feeds, and tons and tons of various tools, and by processing all of that, they provide you with protection.
That's the SOC — it's probably more accurate to call it an organizational unit that receives signals from everywhere and sorts through them. That's if we put it very, very, very roughly. There are caveats with SOC too, as with everything. There are two sides to the coin, the obverse and the reverse. On the reverse side... wait, where did my slide go?
What awaits us with SOC? "Without" should be crossed out, oh well. SOC is expensive, you have to understand that. It's live people plus prepared infrastructure. So, again, we understand we need to calculate the scale at which we need a SOC.
Second point. SOC people burn out often and hard, because it's difficult, heavy work. A SOC has to run 24/7, because otherwise — well, attackers don't sleep. Not on our schedule, unfortunately. The shortage of decent staff is tied to just that. To the fact that, first, there aren't that many great specialists freely available. Second, they get tired fast. That's really true. They burn out fast. Especially if somewhere while building your information security infrastructure you made mistakes, and they're running around all the time, despite SOAR, in fire-brigade mode.
It integrates with practically everything. It all pours into SOC. That's hard too. And attacks keep evolving. So you somehow need to keep raising the level of your analysts.
And what awaits us without a SOC? This. No round-the-clock monitoring, which, as I said, is important. No kind of complete, centralized picture of threats. Response chaos, along the lines of: we've got six infosec guys. Vasya, run over there; Lenya, cut the hoses. And, again, no expertise — not in investigation, but rather for investigations, although the SOC often takes part directly in the investigations themselves. We respond slowly, protect nothing, everything falls apart in our hands. Again, in the case — this is important — in the case of a company scale that warrants it.
So, the takeaways are all fairly clear. A successful SOC is a well-built mechanism made up of good specialists, well-configured systems, and processes.
Our favorite, my favorite — DFIR. Here it's all extremely simple. DFIR doesn't protect anything, but you have to resort to it, because in order to write that same playbook, some new one, in order to understand where it hurts and how it happened, you need to investigate the incident.
So, what can I say about investigation. Unlike SOC, DFIR is a really super deep dive into incidents. It's not monitoring; it very often, in probably 90 percent of cases works retrospectively, i.e. after it's all happened, the incident has already occurred and most likely been contained, but we need to understand what happened.
It uses very deep methodologies, very deep technical tools. So, in a big company, DFIR is like, I don't know, investigating an accident on a nuclear submarine a week later from the wreckage, and while some, I don't know, fleet admiral's yelling at you: yes, you must be done by such-and-such hour. Well, something like that, that's roughly what it looks like. Let's... So, the caveats, what stands in the way of DFIR.
We can't properly investigate an incident, dig into it, if we don't have any systems for recording information. Time is critical; there are data sources used in DFIR that are very time-critical, for example, our beloved RAM.
It takes huge experience, again, not... I mean, it takes huge experience both on the DFIR specialist's side, when he... he needs to not screw up while collecting the data, not wreck anything, not delete anything, and so on. And on the infrastructure side, since, roughly speaking, RAM holds a whole lot of useful stuff, but then some accountant, Valentina Ivanovna, panicked and yanked the cord, everything rebooted, and now we can't find patient zero, we don't understand what happened to us or where it all crept in from. Again, at probably every conference we talk about the myriad of complex DFIR tools, and it's not just the vendor tools that we and our competitors offer. It's also open-source stuff. It's scripts that people knock together for themselves to simplify their own tasks, to implement their own investigation methodology.
And extremely high requirements for data collection and storage. But without DFIR we've become... And DFIR is expensive, it's complicated. I don't know, "expensive" is probably one for you.
But it is quite complicated. DFIR specialists are deep-dive specialists. And the vendor tools aren't simple either. And you need... now a lot of people will say, well, I want my own DFIR team, but I only have, like, three incidents a year that I investigate. Think about whether you need... there are people who provide these services. So maybe it makes sense to contact them, no need to grab it all yourself. Again, assess your scale, do the math.
Without DFIR we don't understand what happened. If we don't understand what happened and how, we can't prevent it in the future. We don't know if the attack is eliminated. Because there really are elaborate, multi-stage attacks. I think that's already been talked about today, and probably will be again. When we have no idea what's going on in an incident, we seem to have put it out, but did we really? Is there no backdoor left, wasn't this a diversion, wasn't this specific thing an attempt to distract our attention somehow?
Again, the absence of digital evidence. I remember about digital evidence, but let's say, if you're a big company and you need to take this into the legal arena, you're going to need evidence. If you say we comply with some specific set of regulations, but we got hacked, and here's what happened, you'll need to show how you got hacked, and that you did everything to keep it from happening.
Without DFIR, again, trouble with regulators and law enforcement. As you know, our reports are read by law enforcement as well. That's also extremely important, because it can greatly simplify your communication. If your DFIR report is done properly, the guys come to you after an attack and say, show us what happened, and you hand them the report. Everything goes faster, less painfully, and with less digging into your business on their part. And, well, increased business risks. Obviously, if we don't investigate, if we go through all of this, or rather, if this whole set of messes has happened to us, then, accordingly, of course, we'll spend more money, we'll spend more time cleaning up the consequences. And the consequences we face will be fatter.
Again, just a thesis. I have a little cheat sheet. Just in case, I think we'll be sending out these presentations. If you want, you can take a photo. This is just a very, very, very basic example, a very simple explanation of what each system is for and at what stage. No, not like that. At each, for the use of each of these systems, you yourselves have to decide in a calculated, weighed, reasonable way, and work out whether you need it or not, and how much you need it. Obviously, if you're under regulators, you need it, there's no way around it, no arguing with it.
Instead of a conclusion, what I'll say. Dima, I'll paraphrase that Telegram post of yours. Don't fall for bare advertising. Approach everything sensibly. Right now it's all colorful, all great, but you need to look, you need to count. Evaluate, consult, communicate with your own business if you're the security team. If you're on the business side, consult your security people. Do we need it or not? What's it going to cost us? What do we get out of it? And, an important point, what's the potential lost profit going to be if we don't adopt this or that technology or tool. And most importantly, just buying and installing cool hardware, just buying great licenses, genuinely great ones.
Right now, security has probably never been better in terms of tooling. We're at the peak now and it's only going to get better. But if you install all of that without proper configuration, you'll just earn yourself a headache, spend a pile of money, get disillusioned with life, and maybe even go bankrupt. Think, reflect, and above all communicate with your business units, otherwise you're toast, gentlemen. I guess that's it. That was my little talk to warm up your brains a bit, because before me and after me there are more strong technical talks. Thank you very much.
Should have ended with: don't fall for bare advertising, but do buy MK Enterprise. Colleagues, your question.
Nik, look, you said very correct things, but even at the bare minimum, what you described comes to those same 5–10 million that small businesses, small and medium businesses, most often don't have. Can you say something about what those who aren't so rich should do? —
— What should the less rich do? There's a very simple, there's a correlation method. I mean, look, there's… How much does hardening cost? —
— No. We could debate this for a long time. In fact, it all comes down to four main areas, based on the threats. That's user identification, password protection or two-factor authentication together with it. That's perimeter control, because everything goes through the perimeter. That's a backup system, given that if it all breaks down, well, you can't protect everything 100%. And that's control of users inside that very perimeter, because users do all kinds of crazy things. And these four areas show very well, essentially, what to do. And here, on user control, for example, inside the perimeter, if we can't control them with DLP, right, well, the average company in the city of Moscow is about 100 people, give or take, revenue around 200 million, and if they have no DLP, and most likely they don't, or they have some cheapest one that's used as access control, then it's logical to put them in a controlled environment, to take something like, if we're talking about budget solutions, Nextcloud, ownCloud, where we can easily control and detect every sneeze.
If our users sit on workstations and again there's no DLP, SIEM, EDR, and most likely there isn't any in a business that small. That's exactly where the forensics story comes in, when you can pull out all the user's actions, well, if he hasn't deleted them, after the fact, yes, of course. And then reconstruct the chain of events that led to a given incident. Well, again, regular backups and, above all, a system for checking those backups, whether they're actually alive or were made a year ago and only halfway. And two-factor authentication, because that's the fight against phishing, that's, essentially, perimeter control. And all of it comes down to a very small kit like that.
It seems like it's just plain simple. At one of the meetups, where... Ah, you weren't there that time. Sasha Dmitriev, unfortunately I don't see him today, said a very correct thing. For each stage he... Just in case, let me explain. Sasha Dmitriev is our great friend, a pentester. Hold on, first explain what the meetups are. —
— We periodically get together as security specialists and give ourselves a brain workout. That's what we call a meetup. And at one of those brain warm-ups we had a lively discussion. Sasha Dmitriev, a seasoned pentester, both physical pentesting and non-physical pentesting. And he said this thing. If you just open the Golden Rules book, open it like this and go through it, everything will be fine. I asked Sasha, I said, Sasha, how many companies have you seen in your whole career that actually went that way? He told me: one. I said, and how many big pentests have you done? He says, 200. With some simple arithmetic we figure out that's only half a percent. You're saying very true things, really, very true things.
But it's just, I didn't ask you that question about hardening for no reason, that same one. I said, how much does it cost? Hardening costs nothing. If you have, well, it costs your specialist's time and money. Now, if you have a good specialist, they can solve a great many problems. But, first of all, users are always the hardest, hardest, hardest part of it.
And what's my point? That yes, at a basic level you can do without certain things, but it will still cost you time. What are infosec tools for? To save your business time, to save your business money. And to protect you from problems that lead to loss of time and money. Can you do that yourself, with some basic methods? Yes, you can. Can you guarantee that those basic methods will actually be followed?
Alas. Let's take one more question. We're running a bit over time. We're sliding into a roast again. I get that it's probably fun, but still. Colleagues, any more questions on the talk itself? On the right. —
— Igor Evgenievich, I see your hand. Nikita, I'd ask a question as someone who deals with clients using Mobile Criminalist Corporate. Well, from our modest experience, maybe three years now of actively deploying it, nobody who's had it installed has dropped it. But the inflow of new clients is lower. I want to ask the following question. The Kaspersky representative who spoke before you was actively trading in fear, he did a wonderful job scaring the audience with what would happen to them, while promising nothing along the lines of: if you install my product, everything will be fine. No, they're not an insurance company, they don't compensate any losses, and neither do you, I suppose. Maybe a talk that's meant for specialists, for discussing the features and nuances of these solutions with Dima, should be reoriented the other way and also start putting pressure on those who make the decisions.
Infosec specialists don't make decisions. What we run into is that they support it, they push it through, but the decision is made at the management level. And that's who has to listen, understand, get scared and make that decision. Thanks. Why don't you do it that way?
— On fear, I'll fix that now. Gentlemen, you're all doomed, our booth's over there, please, without DFIR you're all finished, come to us, buy. Everyone's scared now. Frankly, the people who make decisions, that's exactly why my talk went from highly technical to fairly light and surface-level, because you can go on about tech for a long time, it's heavy and complicated. Scaring people, as if it's some kind of, I don't know, maybe Olga Vasilyevna Pugashkina feeds me this, it's bad form. That is, from a business-ethics standpoint, scaring people is not good. The decision to buy a complex suite — our product is not a simple one.
Plus, it's not so much that the product is complex as that you also need skilled specialists to work with it, who need to be trained. That can be done with us.
But by and large, that decision has to be a weighed one. Otherwise, if we sell through fear now, those people won't come back. As you say, the people who bought our software or any other DFIR software and did it sensibly, they don't drop it, because they understand its weight and importance. If they bought it out of fear, they were scared into it — that's actually what my whole presentation is about. If they bought it out of fear or because someone promised them the moon, well, he took a one-year license, used it or, worst of all, never used it even once. Or the worst thing is he used it and decided he doesn't need it.
What's the likelihood that person will come back later to renew the license? Well, obviously. It's kind of, in my subjective understanding, selling through some kind of fear, or pushiness, or embellishment. No, there's a price, it's expressed in people, it's expressed in time, it's expressed in money. And the person, or decision-maker, either knows about our product and is ready, and finds it necessary, or isn't ready. If they're not ready or don't consider it necessary, we'll wait. The time will come anyway when they'll have to turn to DFIR. Either build their own DFIR team in-house, or bring in DFIR specialists from some specialized organizations that do incident investigation. It's a bad phrase,
"we'll all end up there", but that's really how it goes. —
— Nikita, thank you very much. Let's send Nikita off with a round of applause. Thank you very much, colleagues. And we move on. And overall, Nikita, please hand over the microphone and the clicker. Thank you very much. And I'll put it here. Okay.
15. Maxim Sukhanov (CICADA8) — “The role of reverse engineering in understanding artifacts: when documentation is the enemy and the decompiler is a friend”
Scheduled 13:45–14:15.
Moderator's introduction
You can't always tell, generally, who's your friend and who's your enemy. And friends, like enemies, can be quite unexpected. Our next speaker's friend became the decompiler, and his enemy — documentation. How all of that works, Maxim Sukhanov will tell us. Let's welcome him with applause.
[applause]
Maxim, here's your microphone. The clicker is lagging a bit, let me check. There, take it, please, go ahead.
Talk and Q&A
Hi everyone. I was asked to liven up the audience a bit, so I promise that during the talk I won't say the word "business" even once. When answering a question, I might say it. That's how I'll try to make the talk a bit less tedious.
The topic of the talk, well, you know it; there's a bit about me on the slide, that is, I do cybersecurity incident response, incident investigation, I write various software for incident response and forensics, an NTFS parser, registry, and so on. The problem. The problem is this: how can a person, a forensic expert, as the ideal example, justify some interpretation of some artifacts, traces, or, worse still, anomalies. That is, there's something a person observes when looking at computer data, and how do they explain what happened there. There are basically two approaches: we see some timestamp, and as to what that timestamp is, we refer to an article, a book, documentation, or someone else's results, that is, we say we interpret these bytes this way, and why do we do that, because there are other results out there, someone else's experience.
We can also do some kind of expert experiment and obtain that experience ourselves. That is, the interpretation of some timestamp can be found on our own. I underlined the word "documentation", because from here on we'll specifically look at the example with documentation. In our field there are a number of Telegram chats where people talk, and about a year ago there was a discussion on the topic of this very talk. And there was this small conclusion, which I've shortened further, that if a forensic expert says "I trust the documentation", then in principle that approach can be considered perfectly valid in some situations. Let's see why that's a problem, not a solution.
Vendors, that is, software manufacturers, that is, the very same Microsoft corporation, can deceive people. Deceive, in quotation marks, naturally. It's not deception done out of malice, but an untruth that's communicated for some other reasons. If we open the Microsoft website, the documentation, and look at how they describe the $STANDARD_INFORMATION attribute in NTFS, we'll see that there's no such concept there as timestamps. That is, we see the declaration of the corresponding structure in C, we see the text description, and we see that the places that are used for storing timestamps are reserved. Accordingly, the conclusion: if we blindly trust the documentation, there are no timestamps in the $STANDARD_INFORMATION attribute. Naturally, this is an extreme example. That is, we're trying to refute the original thesis by reducing it to absurdity, but a real absurdity. That is, it's clear here that this can't be the case.
We can also turn to the analogous documentation, the $FILE_NAME attribute, and see that apart from the file name, the file name length and flags there's nothing there, no timestamps at all. There's also a field that references the file record's parent directory. Again, none of this matches reality. It's all deception. But deception in quotation marks, because it has what you'd call a valid reason. The vendor is protecting its internal structures from so-called clumsy hands. That is, it deliberately hides part of the internals, the internals that exist in its operating system, so that people who have no business in there don't poke around. That is, programmer Vasya Pupkin has no business in there, let him get his information from other sources.
And because of that, and apparently only that, the documentation is scrubbed, so the fields you don't need are marked as reserved. This is an extreme example, this is, so to speak, the greatest absurdity one could come across. And I can't, unfortunately, give a comparable one. That is, this is the maximum degree of this kind of mismatch with reality. Also, a vendor that ships some piece of software may forget that its code changed long ago and no longer works that way. And because of that forgetfulness, the documentation was never updated. Here's, let's say, a sore example. There are shadow copies. Ransomware,
when it encrypts data, usually deletes all the shadow copies first. But in some cases, for instance if the encryption was launched over the network, via shares, the shadow copies aren't deleted. And they may survive. If we look inside such a shadow copy, we see that in many cases, starting from Windows 8, instead of user documents, some browser history files and the like, we see a mess of zero bytes. And this effect is seen quite often. That is, people who investigate ransomware attacks know this kind of thing happens.
If you go to the documentation, for the Volume Shadow Copy service, you'll see this shouldn't happen. That is, shadow copies capture all the allocated space of the file system, and you can exclude something user-related from them, but only by setting a special registry value. By default, everything should be included. On the left, an example in a hex editor: from the current state of the file system, a fragment of browser history, Chrome, and below it the same file extracted from the shadow copy. At the top there's data, meaning the file system's current state has data, but the shadow copy has none. That is, the whole space is filled with zero bytes.
And the answer to this riddle is simple. Starting from Windows 8, user files are excluded from the shadow copy. With an asterisk at the end. Since shadow copying is done in blocks with 16-kilobyte granularity, the real kind, based on 1024, while the default cluster size is 4 kilobytes, some overlap with user files can still occur, so part of the user data will still be found in shadow copies anyway. This behavior isn't documented anywhere, and what's, let's say, most wonderful is that Microsoft replied to at least two customers explaining this behavior. That is, they said, in one case, that your data that the ransomware encrypted can't be recovered, because user files aren't included in shadow copies. Thank you, goodbye.
Those replies were found online only once it was clear what to search for. That is, there's a keyword, it's called scoped, scoped shadow copy in full. So you can find a mention of someone citing the vendor's reply in some article.
Right. And this behavior comes as a complete, let's say, surprise to those who were hit. On the one hand, the shadow copies are there, per the documentation user files should be in them, that is, all kinds of Excel documents and so on, but in reality most of them most likely won't be there, if we're talking about Windows 8 and later.
How does this happen? First a full shadow copy is created, but since the copying works in copy-on-write mode, meaning the data isn't actually copied as some backup that constitutes the shadow copy, before it's modified, that is, copying happens as the data changes, that shadow copy doesn't immediately contain everything that will later be shown to the user.
And right after that, the srtasks.exe process starts a rather complex procedure, that walks the entire MFT file, looks for file records, reconstructs their paths and by extension, and not only that, excludes those files in the bitmap of what is subject to copying as it changes. So when user files are modified, the shadow copy won't contain the old version of the changed data. And this applies only to user files, system files, all the EXEs and DLLs. They'll be copied and will end up in the shadow copy in their old state, as required. This is done to reduce the size of shadow copies, since nowadays shadow copies are more a tool for restoring a working, healthy operating system than a backup tool.
On Windows Server versions it's different, and there everything works as it should. So this narrowing, so to speak, of the scope of data copied into the shadow copy doesn't happen there; instead everything works properly, just like Windows 7.
There's also the situation where a vendor documents a file system implementation other than the one it has. If we go back to the 90s and the early 2000s, we'll recall that besides Windows NT there was the 95, 98 and Millennium branch. These operating system branches are different families, different code. Different people, different developers, different code, and whatever overlap there was didn't extend to the file system drivers. And it was exactly then that Microsoft developed its FAT file system specification, which is now used for formatting EFI System Partition file systems. And that specification, unfortunately, is based precisely on the code that was in Windows 95.
Strange as it may seem. So the FAT driver in Windows NT, Windows 10, 11, Windows XP, doesn't follow that specification. And this creates all sorts of situations where developers of third-party FAT file system drivers, say, in Linux, suddenly realize that the requirements which are, as they say today, mandatory requirements, though supposedly requirements are always mandatory, that those requirements aren't followed by the vendor itself, and in particular, because of that, they have to modify the logic, say, of the file system error check, so as not to detect an error where, by the specification, there seems to be one, but in fact there's no error, if you trust the code.
But that's getting a bit nerdy. Another case is a deception that arose from header files being changed on the quiet. A low-level deception, that is. This is Prefetch. Tell me, who knows what is stored in files with names of this format? That is, OP, hyphen, process name, hyphen something, hyphen something else,.pf. A Prefetch file that sits in the Prefetch directory. Well, many probably don't know, because you don't run into it often, and especially not often when there's actually meaningful data in it. I tried, from all over the internet, including the English-language segment, to collect the hypotheses about what might be stored in this file.
Well, there are basically four. The most, let's say, exciting one is the third, because it stems from that, so to speak, unintentional deception, from stale data in header files, while all the others are, so to speak, made up, just hypotheses that people throw out. And the third one is, you could say, not just a hypothesis but a real theory. The problem is that all along, for a very long time, the right answer has been in a Microsoft blog about ASP.NET, but this artifact is only mentioned there in passing, and such a mention in a blog raises the question: is this a mistake? A passing mention, how relevant is it, or maybe it's all correct there, or not. Because the people who write these articles can still make mistakes.
From the blog, the correct answer looks like this. There are certain calls to the Windows API functions OperationStart, OperationEnd, which result in the generation of this file, highlighted in yellow. This is so-called Operation Based Prefetching. It doesn't follow from this, if you just read it, there's no emphasis on it being exactly a file with this name format that should be created. That's the sense that comes through in the text, but no attention is drawn to it being exactly that way and rock solid.
But let's try to take the hardest route, straight from the code. Let's look at where you generally start in such cases. We have a file name template, there's an extension, there's a prefix; we can search for it in the executables inside C:\Windows\System32.
Just those strings. Then filter out the noise. Since Windows ships with debug symbols, we'll know which functions use the strings we find. And the only occurrence left, so to speak, after filtering out everything unneeded, is in the kernel, the file ntoskrnl.exe. Well, or in other files, depending on how that kernel was compiled, but basically it's there.
Hm, the slides are advancing without me. Anyway, we find the name template there, and it's the name without the extension, and something is done with it. And a call to the function PfSnBeginScenario and outside this code branch, PfSnEndProcessTrace. What is this? Where is this function, whose code is shown in the decompiler, called from?
It's called from the function PfSnSetPrefetcherInformation, if a certain parameter with a value of 5 is passed as an argument. And this function is called, the one shown on the previous slide, PfSnOperationProcess. So we have a certain prefetcher, and it has a certain function. If we pass an argument equal to 5 into a call to this function, this argument is called, its field is called PrefetcherInformationClass, then this file gets created, and something is written into it. Question. What does this 5 mean? It's some constant whose meaning is unclear. And if you google all this, there's the Process Hacker project on GitHub. They merge header files from the Windows SDK into it. That is, everything published on the Microsoft developer site is merged in there, along with some community reverse-engineering results.
These particular constants were most likely taken from old header files from Microsoft; a number of signs point to that, including the manner, so to speak, of naming, but that SDK is most likely no longer available, meaning this information came from some old version of the SDK, the developer header file kit, that isn't published now. And the problem is that the constant equal to 5 is called PrefetcherBootControl. That's exactly where it came from: someone looked at where the function that creates a file with this name format is called, looked at what the constant 5 means, and saw that it's called PrefetcherBootControl. And because of that, in various documentation online, not from Microsoft, there's a mention that Prefetch files, OP, hyphen, something,.pf, are created during boot. That it's some kind of boot Prefetch that loads certain pages before they're read.
But the problem is that this name is out of date; in this case the header file doesn't match the code. So this half-documentation, half-something-else is no longer current. It was current once, but now it's all been reworked.
I've switched back a slide. If we move on to the black-box testing method, that is, we know which functions are responsible for creating these files, OperationStart, OperationEnd. If we throw together a Python wrapper for all of this, we'll see that this has nothing to do with booting, with the operating system boot process. That is, these Prefetch files that get created are created simply when the right functions are called, and the boot process doesn't come into it in any way at all. You can pick out, extract certain patterns from the results of black-box testing by looking at the code in a decompiler: first, this trace is created when this function is called, it ends when the OperationEnd function is called, and within a single process only one trace can be created at any one time, and if there are fewer than 32 I/O requests, the results aren't saved, it's considered there's nothing to prefetch.
But to recall how Prefetch works: it looks at which files a program accesses at startup, or after the OperationStart function is called, and preloads those files into memory, so that the data is already in RAM when it needs to be read, so there's no lag from the program reading files very slowly when we need the data much faster. So Windows, on the kernel side, loads that data in advance. So, 32 requests, so-called page faults: if they occur, the trace is saved, and on the next call to OperationStart the kernel loads that data into memory. Straight from these files.
A number of references were mentioned throughout the talk. I honestly cite all of them, so here, under number 4, is the article I managed to find about why shadow copies contain no user files, and it has a link to a response from Microsoft. So, following the investigation of a ransomware incident, people were puzzled by this very question: why is there no user data when shadow copies survived. And they published all of it, but it's hard to find on the internet. You must know what to look for. You have to know the process is called scoped, scoping. So. Various…
Well, that's about it. Anyway, I hope at least some of it made sense. Thank you for your attention. Happy to take questions. —
[applause]
— At least it was honest. He never once said the word "business". So, colleagues, your questions. I think you've broken my audience again. If there are no questions, Maxim, thank you very much.
That was great. Maxim is staying here, so you'll be able to ask him things.
16. Konstantin Titkov (Gazprombank) — “How not to end up needing forensics, and what to do if it can't be avoided”
Scheduled 14:20–14:50.
Moderator's introduction
So, colleagues, let's move on. Usually it goes like this: there's no smoke without fire. But in fact, once it burns, it's too late to put it out. In general, with information security the situation is about the same. So, how not to let it get to a fire, and what to do if it happens anyway? Our next speaker, Konstantin Titkov, will cover that. Let's give him a round of applause.
Konstantin, here's the clicker. Advance the slides over here. And the microphone is on. —
Talk and Q&A
Thank you. Hello, colleagues. Our organizers sensibly alternate talks of different kinds. I'm essentially continuing, logically, the talk Nikita Vyugin gave earlier. I'll give you a lot of tools that may come in handy when talking to your management, to business owners, to clients. So please, don't hesitate to take photos, take notes, move up closer, unless of course you've got a super telephoto lens. Otherwise you may miss a lot. At the end, links to materials with QR codes. I recommend photographing them and taking them along. So, we're going to talk about what to do if all those preparatory measures, EDR, antivirus, firewalling, didn't deliver the expected result, and your cybersecurity risk has materialized: your infrastructure is encrypted, the data stolen, or everything just wiped.
What do we do about it? A word about me. I head the cybersecurity center for Gazprombank's subsidiaries. I'm also an ambassador for Cyberdom, meaning I help get important cybersecurity information across to Russian business. What's on the agenda today? What's wrong with our typical IT incident response plans? What emergency actions can be taken? Who can help, and how? Spoiler: a company that's genuinely good at DFIR, of course. The stages of incident response, what to do afterwards, the specifics of a personal data leak, and above all, what a business can do in advance to recover faster and make it hurt less.
A few examples that happened literally in July of this year. A large federal beverage producer, a retail chain. Hit by ransomware, shipments stopped, sales stopped. What did customers do? Customers went to another store round the corner, bought it all. There they got a loyalty card and were told: come again, we'll sell you more. Same thing with a federal pharmacy chain. Direct loss of customers. That's only the tip of the iceberg. A major airline was recovering. Two weeks later I read the news. We are still calculating the amount of fuel for refueling our aircraft not with the help of specialized software, but based on fuel consumption statistics for flights on similar routes over recent years. So what can we take from this? Not all of the business could be restored even two weeks later.
Moving on. So what does a business that's been encrypted or wiped try to do first? It tries to recover, to restore its core processes. And in doing so they forget to find the attacker: how he got in, all his persistence points, tunnels, backdoors, web shells, compromised remote-access accounts, newly created remote-access accounts disguised as service accounts, and so on. So they try to restore something, the attackers watch this, delete it, destroy it, see that the efforts are failing, and raise the ransom. Well, if they're financially motivated. Whereas first of all, of course, you do need to kick the attacker off the infrastructure first and only then fix and restore it.
So, an IT incident. What do we usually have? A software failure. One-off or recurring. A hardware failure. Human error. A screw-up, that is. Someone entered it wrong, deleted it, broke it. We restore from backup, replace the hardware, roll back the software version, contact tech support. We have staff for this, because these are expected incidents, we expect them to happen. We train employees for this, we buy technical support for this, we have operating manuals and so on, even some kind of recovery plans. Even backups. Maybe there are even some among you who have at least once tried restoring backups and checking that it works.
If you haven't done that, please do, it's extremely important. Like that story with GitHub, when they went down, they had six backups and couldn't bring any of them up. Okay, what is a security incident? It's a deliberate malicious action against your infrastructure, the whole of it and any element of it, against your personnel, your employees, and it's adaptive, taking into account how you try to resist, and restore it. And if the correct action in this situation is to work from a pre-developed cyber incident response plan, which is drawn up on the assumption that we come to work on Monday, and all we have available is the phone in our pocket and all that was switched off, and all that was switched on, consider it unavailable.
What will you do, and how long will it take to restore the whole infrastructure, data and processes? Well, actually, in practice, it usually starts with time lost to panic. The IT department tries changing passwords, tries restoring something from backups, the drives with backup copies get connected to the infected infrastructure, the attackers, of course, wait for this, and those backup copies on those connected drives, what? They delete them. Great. And an hour, two, three, a day later, I know cases where on day three it dawns on us that we're doing something wrong, that we've wasted our time, effort and money, and we must respond to the incident, that is, do DFIR, Digital Forensics and Incident Response, rather than just trying to fight it as an IT failure.
Well, when does it usually happen? Friday, public holidays, weekends; lately it's started happening on Thursdays sometimes. Why? Because ransomware or a wiper needs time to destroy the whole infrastructure. They're write operations, they take time, so ideally you don't get in their way. And here it's vital to have employees' phone numbers to call them in to work, recall them from the dacha, a picnic, vacation or a business trip. You do have the numbers of employees' relatives, so you can reach them on a day off, printed out on paper, right? And you have to rely on all the groundwork the company has done in advance. Because only that will let you (a) recover, and (b) do it as fast as possible.
And here, of course, the ability to call in help is invaluable. Why do I draw that conclusion? Because I've spent several Fridays on the phone talking to my friends, who came to me for help and advice when the infrastructure they were responsible for had been encrypted. And we were discussing, you understand, not old classmates, but the sequence of steps for restoring their infrastructure. And after two or three such conversations I concluded that this is probably not how to spend a Friday. My colleagues shown on the slide and I drew up recommendations, because, unfortunately, we couldn't find any publicly available recommendations online that were sufficiently complete, correct and exhaustive. And so today I'm presenting them to you, I'll run through them briefly, with links at the end to download them for use.
The purpose of the recommendations is to give an understanding of what to do, what needs to be done in advance so that you're able to do it, and an understanding that you have to act, and act fast. We describe how it could've happened, the main intrusion vectors, so that you can convey to the client, to the business owner, to the CEO, what exactly, if you don't deal with it, will increase the likelihood of this sad event. We consider whether to pay the ransom. In a nutshell, we remember that ransomware is a program, and if you try to buy the decryption keys from the attackers, first of all, they may not work, because what matters to the attackers is encrypting your infrastructure, while decrypting is a secondary task for them, they may not have tested it, that's one thing. Secondly, you may simply not get the encryption keys for your money, especially being in the.ru zone, a Russian company and so on.
And you have to understand that your bank may not just fail to help you but refuse to process a payment to the attackers, citing anti-money-laundering and counter-terrorist-financing rules. And worst of all, if you do manage to push that payment through, it may later be classified precisely as financing of terrorism, with all the sad consequences for everyone who took part in making that payment. So, the main emergency actions, and this is only the tip of the iceberg. Check online banking transactions. Because of course your accountant has a printout with the addresses, phone numbers of whom to call, with what details, to verify the payment orders that were sent but perhaps not yet executed. Maybe there are malicious ones that can still be cancelled and the money recovered.
Don't power off or reboot the equipment, to preserve the forensic data in memory. Isolate what's valuable, above all the backup system and the backups. Take snapshots of the VMs, if still possible, the virtualization system isn't destroyed. Examine outbound traffic, to see what was downloaded from you, stolen, and assume a possible data leak. Consider network isolation of the company as a whole or of certain segments. Take care of the staff's work schedule, since the company is now switching to 24/7 mode, most likely for several days, or maybe weeks. HR — take care of how they'll pay for the overtime, how they'll process it. PR — how they'll interact with the regulator, with clients and with the press, preferably with press release drafts prepared in advance.
IT — where they'll restore from and which contractors they'll bring in to help as extra pairs of sysadmin hands and onto what hardware or into which cloud they'll do that, Security — which DFIR expert firm they'll bring in, and so on. The roles are spelled out there too. Each can carve out their own piece of the recommendations and do their preparatory homework in advance. As for help, yes, of course, we mean companies that are good at DFIR. There are recommendations on choosing the company and the engagement mode. Obviously you agree on the number of experts, date and time of the visit, address, contacts, meeting, passes, parking, kick-off meeting, planning. What do we want?
Do we, say, want not just to investigate but also to prosecute? Or do we want to investigate and, if possible, avoid publicity? Prosecute without publicity? Those are hard to combine with each other. The experts will point that out. Analysis of the compromised hosts, identifying all the entry points, disk analysis, and so on, and so on, and so on, clean-up, and only then do you get on with recovery. Obviously you're interested in restoring the company as quickly as possible, we have to save time, you need to know how many DFIR experts will be working, and it's important that you can provide the same synchronous working mode for your own specialists.
Otherwise they'll sit and wait for you to wake up and come to work, and they'll be idle. Need I say that all CII entities must first of all report to GosSOPKA or to an accredited GosSOPKA centre, or to the NKTsKI. What to prepare for the expert's arrival? That's a workplace, and communications outside the affected infrastructure. Because otherwise the attackers will be reading your email, your messenger, your personal messengers, if you logged into them from equipment connected to your infrastructure, they'll read them and take measures to obstruct you, knowing what you're preparing to do. Moreover, they'll post screenshots of your correspondence online and mock you in every possible way. So it's better to set up communications outside the infrastructure
at once. We talked about synchronous mode. Cancel vacations, business trips, resignations. That is, everyone who can work must be called to arms. Now let's run through the response stages. Many variants here, based on the SANS guide. But I'd like to point out that the first one is incident preparation. What you as a company, your clients, can do in advance. OS, application, DBMS and security tool logs. Do they exist at all? Are they detailed enough? For how long are they retained, and where? Won't they be wiped or erased or encrypted together with the infrastructure? Disks. Physically, a DFIR expert can find, together with you, exactly where those hard drives are, in which server.
That is, map an IP address to a specific rack unit. Pick them up, physically unscrew them, make sure there are no locks. And if you used, say, your own protection tools like BitLocker, then decrypt them too. Artifacts. Can you run the necessary program on every host in your infrastructure? Do you have network access? And access rights, an account, and do you remember the password for it? And it's surely not written in a file on the server that's going to be encrypted? Such cases, unfortunately, are unheard of. Reserves. Do you have a financial reserve for bringing in a contractor, for computing capacity that you'll be restoring onto, since some of your equipment may later be subject to seizure?
Backups and so on. And have you backed up the installation packages? And the license keys for those packages, have you backed those up too? Especially for "trophy" software, where you can't request them from the vendor again.
Stage two is identification. It's great if you use TI feeds. You may learn of a planned attack on your infrastructure before anything has happened. A kind of shift-left. Then perhaps your whole response comes down to two or three hours, and there won't really be any downtime. A compromise assessment will be enough. If not, worse. Perhaps your SOC will react and say something suspicious is going on. If you don't have one, well, maybe you at least have managed EDR, and the partners will say, guys, something's suspicious on those hosts. No managed, then maybe at least your own EDR. Well, it's worse when we just see in IT monitoring antiviruses being methodically disabled on hosts. Also suspicious, you'll agree, if antiviruses get switched off. And attackers will disable security tools, they simply get in the way. They might block, say, the ransomware launch, or whatever. Well, worst case, if you haven't taken care of the rest, you'll read in a Telegram channel on vacation that there's no point coming back, nowhere to come back to.
Containment. That's probably the last thing a company can do on its own, without any help. Isolate the equipment, don't reboot, but disconnect from the LAN. VMs, physical machines, other interfaces too, and disconnect from the SAN. Particularly cunning and nasty malware can, for example, trigger encryption if it loses contact with its command centre. That is, you've cut your infrastructure off the internet so that nobody interferes with your investigation and recovery. And some ransomware launches precisely at that moment. So, accordingly, if you have a mature infrastructure, all your data is on storage arrays, accessible via the SAN, the data network. You disconnect that network, and whatever's running in the OS, the ransomware, has nothing to encrypt. All on disconnected storage.
Cleanup. This is exactly where you need experts, of course. That's essentially finding and removing every single backdoor, no exceptions. Restoring security controls that may have been disabled, tampered with, or had exclusions added to them. And then changing passwords. Everywhere. Domain, local, on *nix boxes, Macs, Windows, standalone workstations, in DBMSs, VPNs, applications, external online accounts, access certificates, one-way and two-way, the tax service (FNS) portal accounts, and so on. It's crucial, I'll explain why a bit later, not to forget anything at all. And don't forget the system accounts either.
Well, recovery and lessons learned. This is where you restore all your data from backups. And most likely you'll first set up the apps from the installation packages, applying all the patches, and only then restoring the data from backups. If you still have them, of course. For example, if you use the 3-2-1 backup scheme. Three backup copies on two media, one of which is outside your infrastructure. And preferably with integrity checking, too. You've recovered, you've restored the data. Those without backups start looking for test environments. Listen, I think on the test environment we once deployed production data for testing. Let's restore from that.
Bad idea, of course, production data in test, but everyone does it differently. And some say, well, so now, in three shifts, the whole team will key data into the database from the paper originals. Two weeks or so. That happens too. Incident response report. Recommendations, a plan of work so it doesn't happen a second time. All this can take a while. I mean, as I said, if the ransomware hasn't launched yet, maybe in 2-3 hours you'll sort out the whole situation and kick the bad guys out. But if your IT team now goes and mass-changes passwords, then the attacker goes, oh, they're changing passwords. They've probably noticed me.
Well, we launch. Godspeed. And we go straight on to the second one. Encryption has run, and here it's already measured in days, weeks. Well, if you prepared, then days; if you didn't, then weeks. It can be done online, remotely. Especially relevant for those whose infrastructure is far from the federal centres. It just takes the experts ages to get there, 3 days by dog sled. But here it's very important that the hands on site will be yours, your staff. If you don't have enough people or enough expertise, say they're all juniors, that happens, then it may end up only worse and take longer than if the experts came to you.
And unfortunately that's not all, plenty of fun awaits you afterwards. That is, you go to law enforcement with the complaint materials, the DFIR report that you prepared either yourselves or together with the experts you brought in, crime report register (KUSP) entry, getting a number, then possibly a seizure, opening a criminal case, a certificate recognizing the company as victim, so that when dealing with other government bodies, say, if you got encrypted at the end of a tax or reporting period, you're able to say: we're not hiding our financial statements from you, we just can't submit them on time, don't execute us.
But that's not all either. They may have stolen data, for example, personal data. And if you paid the ransom earlier, for example, then it's a nice move to ask you for a second ransom for deleting that data. But as a rule, nobody deletes it. Even if they got the ransom, it's still sold, published later. And then roughly 6 billion people, give or take, the adult population of our planet, can do whatever they want with that data. What will they do? Now, for the personal data, the regulator comes straight to you. Customers, obviously, will be unhappy and will file complaints. If there's customer data in there with the contract terms, then your competitors go, oh, good customers, they pay on time, let's give them a small discount and poach them. Interesting?
Your intellectual property, for example, source code. Great, convenient. Keys and passwords. This is exactly the big problem: roughly several billion people can now calmly parse your data, and everything they find there, logins and passwords, try to use in your interface or in those of your counterparties, customers, and government regulators, to log in as you and do something nasty. So back at that stage we changed everything. If personal data leaks, the notifications. And the most fun part. Be prepared that in future, 100% and more than once, maybe even every day, information will appear on the internet that you've supposedly been hacked again and your data stolen again. And sample data will be published, it's your old data, plus enriched with made-up, generated data or data from other leaks. And now every time you'll have to go proving your innocence.
You'll be comparing this data posted on the internet with what you have now and what you had at the time of the incident, and trying to work out if we've been hacked again. Or if attackers are just trying to pass off a compilation as a new breach. So either we say it's all a lie, or we say we need to do the DFIR all over again. And if it isn't all a lie, then in fact in every real case you must notify Roskomnadzor within 24 hours, including about a fake leak. You still have to notify within 24 hours, and within 72 hours send a report on the internal investigation's results. And we said that DFIR takes several days even if prepared. It's not some other 72 hours.
CII entities must notify NKTsKI, entities under the Central Bank, FinCERT. Well, maybe someone also complies with GDPR, I don't know. So for every body you have to inform, it's advisable to figure out in advance who will do it, based on what template, where to submit the report, who to coordinate with, who they are, whether they're authorized, and so on. And if the interface is unavailable, through your fault or not, how do you deliver it on paper and where, and how do you get a receipt stamp. Well, accordingly, my personal recommendation for preparing any report is to make it complete and exhaustive, because government regulators have little time either, and getting into correspondence isn't worth it. Better to describe in full right away what happened, what's been done, what we plan to do.
Well, accordingly, if the worst hasn't happened yet, but there's a suspicion it might, you can carry out a compromise assessment. It's a sort of light version of DFIR: logs are collected, traffic is collected and a probabilistic conclusion is made: has your infrastructure been compromised, the attackers are already inside, or not compromised yet. Why? From the reports of the largest cybersecurity companies and the reports of digital forensics labs, we see that before launching encryption or a wipe in the infrastructure, attackers may spend 3, 6, or even 9 months there, doing reconnaissance, downloading all they can, looking for backups to delete them too, or encrypt them, so it's harder for you to get back up.
And it may be that during that not-so-short period, if you get suspicious, you run a CA and find yes, the attackers are there, they're preparing to destroy your company, and with far fewer people and resources it'll be much cheaper to counter them and throw them out. How to prepare? CISOs, CIOs, CTOs, directors of IT and cybersecurity can develop in advance a response plan for exactly the case of a successful cyberattack with the most dire consequences. And it's not an IT recovery plan at all, it's a completely different kind of plan. Create the necessary reserves for this case, including finances, hardware, software, and a lot of printed paper, in advance.
Pay special attention to the security not only of data backups, but also of backups of application software and license keys, and protect the backup system itself. Because this story, where the backup system, its server sits in a flat network with the hosts it backs up, and the workstation from which the backup system is administered is also in that flat network, and domain-joined as well. So, do you think you'll have any backups left if ransomware runs? No, of course not. And, actually, run response exercises, even if only tabletop ones. We came in on Monday, nothing works except what was powered off.
How much time will we need to restore operations fully or to whatever extent the business needs? And is the business okay with that timeframe? Is it fine? If not, investment is needed. And there are two kinds of investment. Investment in reducing the likelihood of an incident. That's antivirus, firewalls, SOC, EDR, and so on. And then there are investments, reducing downtime and reducing the damage when an incident does happen. Those are different measures. And some of them, just as with hardening, you can do in advance, calmly, for free. Not all of them, but a great many. Including choosing your DFIR partners in advance, or running a compromise assessment periodically, or something else.
How can a CEO prepare for this? First, pass the recommendations to your IT and security teams, check, make sure that everyone can do it all, knows how, and nobody's embarrassed to say they can't do something or lack something. If needed, give them what they need. Estimate the operational downtime of the company in an incident with the grimmest consequences. And decide, honestly, am I OK with that downtime? Are these losses, reputational damage, customer churn, fines, inspections and so on, acceptable to me or not. During an incident, bring in every possible resource, support with everything you've got, support the team, support the employees. And the main thing to remember: any employee who let an incident happen, through their own fault or not, but who is demotivated, will have no interest in cleaning it up. But the one who took all measures and every effort to get the company back on its feet, to eliminate the consequences for the company, is worth two new ones.
Recommendations linked, they're published together with Cyberdom, formatted, available for download. Please use them for the good health of your companies, your partners, your clients. Thank you for your attention.
Konstantin, thank you very much. Questions? Thank you, very interesting. Here's my question. Isn't it a bit late to shut the barn door after the incident? So, from experience: cases where encryption, including of the backups, has already been done, versus cases where it was caught only at the very start, that is, an incident still in preparation versus one already carried out, Which is there more of, as a rule? You see, it would be more correct to put that question to the companies that get called in to do DFIR. Because you and I can look at it from the outside, but survivorship bias is a thing. Many of those who couldn't recover don't come to these conferences don't talk about it. Because the business is completely destroyed. Well, sure, so they hide the information, but we all see it once it's already happened, but overall...
Our job, yours and mine, is to prepare so that when, or if, this happens, a) we recover, b) we do it as fast as possible. —
— Well, that's clear. Thank you. And a lot can be done in advance, easily and for free. Start by printing out what will be unrecoverable in case of encryption but will be needed first. I basically said that at the very beginning too, so thank you. —
— So, colleagues, more questions?
Konstantin, I see no hands, so thank you very much. Let's see Konstantin off with a round of applause. Right, the microphone. And we, colleagues, are now going on a break, and we meet again in this hall at 3:30. Thank you.
Thank you. On to a separate block of today's conference. So I ask everyone to please return to the hall. You can bring along whatever you picked up in the catering area. And in literally a couple of minutes we'll start moving on. Thank you.
[music]
[music]
[music]
[A break (14:55–15:30 in the program) is cut from the recording; about a minute and a half of music remains from it.]
17. Nikita Pavlov (Zero eDiscovery) — “The role of the user interface in digital investigations: how UI convenience and intuitiveness speed up data analysis and search”
Scheduled 15:35–16:05.
Moderator's introduction
Testing, testing, testing. So, friends, as I already said, let's open the third and final block of our two-day conference. What is the role of the user interface in a digital investigation? And how does all that convenience help us cope with it? Nikita Pavlov from Zero eDiscovery will tell us about that. Let's give him a round of applause.
Next slide. The clicker. Flip it over here. And the microphone.
Talk and Q&A
Good afternoon, dear participants. Last year, those who were here probably remember, we talked about a methodology for conducting investigations using forensic data analysis systems. This year we decided to diversify our talks a bit, I suppose, to diversify what we narrate, what we talk about. And UI turned out to be the topic that's probably covered the least, in the work of law enforcement agencies, in the work of people who do forensics, computer forensics, various private cases. Accordingly, as co-founder and co-creator of the Zero eDiscovery platform, designed for conducting various investigations and audits, and the actual designer of its user interface and user experience, I'd like to share some practical examples, practical tips on how you can improve your interaction, first of all, with the developers of the products you use, and secondly, I suppose, highlight some interesting features that you may not have noticed before.
The first thing I want to say, to raise as a problem, is one that is present in most interfaces, basically, that are developed by people who aren't specialists, not by people who are actually trained, who went through a certain school of training in design, in engineering. The word "design" in itself gets lost, and it's usually understood as a way to make something beautiful and attractive, but at the same time, I think the word "design" should be understood the way it's understood in English and used very often to mean engineering rather than decoration or anything else. So the first thing I want to say is information overload. In almost every interface I personally encountered during my work, which I at some point rewrote or whatever else, very often everything is piled into one heap. And you've surely come across situations where you're in an interface and just don't always understand how to accomplish a particular task that you're facing.
Accordingly, there's an interesting phrase here, Hick's law: the more choices, the longer the decision. Absolutely obvious: the more buttons in an interface, the harder it is to work in it, because you have to hunt for them. Not if, but surely everyone uses Microsoft products: Word, Excel. Some have probably switched now to Yandex or VK products. And everyone knows perfectly well that in those interfaces you use at most 6 or 7 functions. Accordingly, my goal at Zero eDiscovery, in development generally, is to provide the clearest possible interface, as simple, intuitive and accessible as possible for accomplishing the primary task, rather than learning a new software product.
Again, I'll repeat what I said at the start and expand on it a bit: user interface and user experience are not about beauty. Although there's something to the aesthetic side, but, as perhaps many will confirm, many have come across this: in the world of, say, the military industry, everything is very beautiful, certain objects, for example tanks, maybe combat vehicles, aircraft and so on. And that beauty we see is directly tied to the fact that, first of all, nothing's superfluous, because everyone knows perfectly well that the more loaded up a certain machine, a certain device is, the harder it is to control, and accordingly the harder it is to operate. So the first term we want to remember today is what user experience actually is in the first place.
User experience is also comparable to the word ergonomics, which you've come across and may have used for objects of the physical world. Ergonomics in itself is a property of an object, first of all the fit of its external parameters to its purpose, that is, a hammer, say, that's comfortable to hold, a screwdriver that is comfortable to hold, an ergonomic handle is used. The same can be said about a tactical grip. That is, all of this is first of all not about aesthetics, but specifically about utility. Accordingly, user experience is, first of all, easiest to understand as the utility of the software product itself.
User interface, in turn, is its software implementation, its implementation in the interface, that can be considered the part that is written in code, designed on some other platforms, which are still code anyway. Accordingly, the goal of user experience, and user interface is, basically, to increase the productivity of a given employee when performing some task. That is, its job specifically is to give you convenience, clarity, and help you get some piece of work done as fast as possible. Another aspect that goes a bit deeper is cognitive load.
When we talk about cognitive load, we're talking about our so-called working memory. Every person perceives, and everyone has heard this, 7 plus or minus 2 objects at a time. Surely many of you have already encountered this in the context of other subject areas. In our case, same thing. When we look with our eyes, we're not able to take in too many objects, we're not able to take in too many colours, not able to take in some unlimited set of functions, some super-capable thing and so on. So the most-used products you could name right now, and you'll surely agree with me, is Google itself, which has no functions whatsoever compared with, say, Mail.ru, which has probably been around since...
By Mail.ru I mean the search engine, which has probably never changed at all since 2005 or so. That is, overload doesn't let you use the search engine effectively, so everyone usually goes either to Google or, these days, accordingly, everyone already works in Yandex. Yandex drew on the experience, the long experience of all search engines, and put it all more or less in order and provides a really quite convenient interface.
Accordingly, our job when implementing an interface, and your job when using it, is to find the roughest spots and eliminate them to increase productivity. Now for something a bit more interesting. Gestalt principles are probably not the most common terminology that you come across in relation to user interface and design. These are basically the fundamental principles your understanding is built on, your work, even paperwork. So if we talk about the principle of proximity, about elements placed next to each other, they are perceived as logically related. In the same way, if there's a computer on your desk, say, and a cup with coffee in it, you'll understand that these objects are related. Though logically you'll understand they shouldn't be there together, because that violates safety rules.
Accordingly, in an interface, when we work, we always need to take the related components and place them in one area. Here I'll explain, a little digression, why I'm talking about these. Because every time you interact with an interface from now on, it'll be very interesting to apply each of these principles and write to the developer: listen, please change this to that, because these things aren't logically related. The principle of similarity. Colour, shape, whether it's a geometric shape or an organic shape. Many of you've seen this, for example, in the interface of any email client.
What's related there? On the left you have the list of messages, on the right the message view. That's exactly what we mean: separate areas are placed in different locations. Likewise, in the tools you use at work, you'll often find there can be, for example, dedicated pages for user settings, dedicated pages for, say, interface settings. Dedicated pages for configuring some other parameters. If you're loading data, a pop-up window will definitely appear, and so on. So through shape, colour and various contrasts we can provide a much clearer understanding of what we're doing right now in our work.
Next, regarding the principle of closure. I'll skip it here, it's a bit convoluted, and I'll come back to it in the Q&A section. As for the principle of common fate, the point here is that in every interface everything must happen in a predictable way. There's a, so to speak, a phrase from a certain book: "this button was here the last time I was here." Some of you may have noticed that in web apps, modern ones, the placement of certain components changes very often, and accordingly the layout may change after updates. So after an update you usually don't immediately understand how to perform a particular task. Especially in my case, probably, given my experience in the organisations where I've been, I worked more with American software. Take Google Workspace, it's a set of tools for email, presentations, documents and everything else.
When, probably in 2018, they started reworking the workspace they originally had, and around 2020 they released the update, literally everyone, I think, was lost and had no idea whether to keep using it, because the changes were too, were too significant. Yes, so here we're saying that every element should still be preserved: if we've already laid down some particular pattern, that is, a user experience, we absolutely must keep seeing it. And the figure-ground principle is much the same as similarity and proximity, just more, just deeper, let's say.
As for, for example, how a user interface can, there's a lot of text here and little meaning, I'll expand on it, how actually a user interface can effectively help perform a specific task. If we, for example, during any incident, we open the logs and, roughly speaking, there's a wall of white lines on black, maybe even without any highlighting, our eyes run over it, we have to read 10,000 lines and, accordingly, well, maybe by the hundredth line our cognitive load, which we talked about, is just so, the brain gets so overloaded with information that you stop, accordingly, taking in any further lines of those same logs, or if you're analysing some tabular data. Now, if instead of that we had an interface that simply indexed at least those logs, and we could specify some concrete points to search for, we would basically cut the work down by tens, if not hundreds of times.
So, the second element, what we talked about, figure-ground. If we, accordingly, highlight things in the text with separate colours, red means alerts. If we've taken some pattern and, through rules, programmed it, green is good, orange is moderate, needs attention, red, accordingly, is highlighted, and we see that something is bad. So when we talk about the actual benefit of a user interface, we're talking about how conveniently placed components, conveniently placed elements, properly in the right hierarchy, allow you to take in information much faster. Accordingly, in real life you may encounter a similar, let's say, pattern, if we're talking about a stack of papers. If you bring, say, your boss a stack of papers like this, which sheet will he read first? Accordingly, you need to put the right sheet, the one that you and he, accordingly, care about most, on top.
Same thing here. The interface should highlight the key aspects. The interface should highlight the most important details. And now, what I want to say about practical points. This slide will probably be worth photographing for later. But overall, when you now interact with the developers of any product that you encounter, that you work with, be sure to try to pay attention to whether the interface is convenient for you, whether it's actually effective, whether it lets you get the job done. And you can pay attention to each of the points I've listed here. First, regarding unified search. If the platform some vendor has built does have search, ask them to make it a single one. Because the more search engines inside one platform, the harder it is to work.
Second, regarding templates and dashboards. Nowadays you absolutely have to build interfaces that let you customize, basically, your work. Because even if you look at your own desktop, at your colleagues' desktops, one likes it here, another likes it there, someone likes two monitors, someone doesn't work with monitors at all. And accordingly, it's the same here. Demand it, ask for it. It's interesting, and it really does improve work efficiency. And be sure to ask for visualization. Statistics, various charts, various links, if you work with some kind of structures, if you work on investigations, if it's searching for affiliated persons or for relationships, some kind of corruption scheme.
Hotkeys are probably not the most common way of getting work done in general, but those of you who have used Excel, who use Excel actively, for example, may know about the Alt key, which activates shortcuts that let you perform basic functions much faster, for example sorting data, arranging columns, their sizes, enlarging, shrinking and so on. Knowing them, you simply multiply the efficiency of your work many times over. Clarity and unambiguity — again, it's the Gestalt principles that came up earlier; it's about everything having to be clear. If a button is red, it's most likely delete. If a button is green, it's most likely create. If a button is blue, it's most likely some more or less neutral action or running some operation. Try to pay attention to that too, and don't hesitate to point it out to the developers.
And as for feedback, absolutely demand that the system talks to you. If you're working with a software product, you click a button and it complains sort of on the sly, and you don't understand what's going on, be sure to write to the developers saying, listen, it would be great if, when I click this button, it at least showed a spinner. So here I want to stress: formulate and voice your requirements for the interfaces you use. In terms of benefits for business, and for working in agencies in general, a good interface definitely lets you invest more time in the truly key tasks.
Surely everyone, or many of you, those who have worked with various technical tools, may remember the evolution from the 2000s to around 2015, when really all the STS, the special technical means — for me the closest example is customs control.
Technology is developing very fast and, accordingly, raises labour efficiency, raises the efficiency of individual employees, lets you focus on more interesting and important tasks, the ones our computers and devices can't do yet.
Summing up, I'll repeat once more what I want to stress: that users pay more attention to how things should be, that users get involved in this part and remember that you can always get feedback, and try not to put up with people building bad, inconvenient interfaces. That's all. It would be great to discuss any questions, if any come up. Colleagues, your questions. Thank you. —
— Right, Alexey. Well, I'll ask one question, as a direct user — not of your interface, but of a rather scary and overloaded one; the cobbler's children go barefoot. So, it was strange not to see — I think I didn't see — an item like customization. Because even a bad, even a clumsily drawn interface — to hell with it, so my button is red, I'll get used to it. It's another matter when I need, for example, the filters arranged in a specific order. To put the ones I need there, not the ones the developer thought most popular, and the column order I want while investigating something in that interface. Again, in 99% of cases some four main columns are enough for me, and I want to hide the rest, or for different data types — a separate set of columns for email, and so on and so forth. So drawing little buttons is great, of course, but unless you're a client with tens of millions a year in support fees, you can only change the colour of the buttons in your imagination, and the developer is unlikely to adapt to you.
Yes, that's a very good observation. I'll say two things about it. First. I probably didn't elaborate enough on the templates and dashboards point. That's exactly where I wanted to talk about the ability to adapt the view to the tasks you need. This is precisely about the interface letting the user customize something for themselves on their own. But the second thing I want to say concerns the possibility of customization. Why I'm not exactly against it, but I won't champion this particular idea — because other people may be using the system besides you. Accordingly, one way or another, the user experience, the actual user experience, should be close to uniform, because you have colleagues whose colours and layouts may, accordingly, be different. And if, say, you recolour the buttons — well, buttons are the most basic example, but if we're talking about columns: you set them up your way, did the task, colleagues came and said, look at the second column.
His second column is completely different. So with customization, especially in products that are used collaboratively, you need to be very careful, otherwise there's a chance of catching a caveat, so to speak, where problems arise in your communication simply from misunderstanding, if you've really over-customized things in different directions. But overall it's a very valid, very fair remark that the user interface should be highly customizable for the individual. And that's what I wanted to say here. Let's say, in incident investigation, in this respect — if we focus specifically on this — I think this problem is solved by an end-to-end identifier for a piece of intercepted data or whatever, some ID that is simply consistent throughout. So which column is where is basically not a question. And on top of that, it seems to me that generally there are about 10 people on an interface...
Oh, not on an interface — on a single incident, some phishing email, there aren't 10 people poring over that email at once. So we'll stock up out there and continue the argument with knives. —
— Fine, just please, no blood. At least not today. Any more questions? Anyone? Well, if there are no questions, let's send Nikita off with applause. Nikita, thanks so much. As always, it was great. The mic — and that can stay on the stand. Thank you very much.
18. Alexey Drozd (SearchInform) — “Using steganography in various channels to identify the source of a data leak”
Scheduled 16:10–16:40.
Moderator's introduction
Well, as we've all figured out today, sometimes a single file, some big one, can turn out to be something other than what it seems to be. So, for example, some holiday photo may hold not just palm trees, but also, say, an Excel file. That's what's called steganography, which today Alexey Drozd will be telling us about. A representative of SearchInform. Let's give him a round of applause and see what we've actually got here.
Talk and Q&A
Hello everyone. I've only let go of the microphone for a short while, it turns out. I wanted to rework the talk, but it was too late. Rework it how? To add a callback. Those who were here last year or remember last year's talk, I spoke about methodology, about the incident lifeline. So where does steganography come in? I look at it from one particular angle here, although users themselves also gravitate towards steganography to exfiltrate information, but more often, since you can't install just any software, or since steganography is still not for average minds, writing your own engine for it and getting further than copy /b and file concatenation.
So users, for now, still mostly prefer cryptography, but in the sense of very primitive encoding, that is, attacking what DLP systems, for example, or other monitoring systems can't recognize. The most basic thing, which analysts almost the world over are waiting for, is that large language models will arrive, including LLMs, and solve this kind of primitive trick. A user exfiltrating some personal data hits Ctrl+F, find and replace, and replaces every digit 5 with the word "five". So all our personal data, which used to consist of, say, digits and was caught perfectly by regular expressions, now isn't caught so well by regular expressions. And there are a gazillion ways to come up with a trick that none of the DLP systems I know will see through without some neural nets and so on.
So, steganography is also something a security officer needs. What does he need it for? Right, aha, now I understand why people kept switching the wrong way. Here's the UI/UX interface for you. The top button goes back, the bottom button goes forward. There you go, now you know. So, why does a security officer need steganography in itself? Precisely in order to lay down some straw in advance. If you look up or recall the incident lifeline, what is it? It's a timeline, and on it there's a point of no return, when the leak happened, when the information was taken out. We, as the now former owners, don't know what's done with it, how it's modified, where it's being distributed, who has access to it. So, the incident has already happened.
This gives rise, in the minds of the security department, to a stage of an acute sense of impending apocalypse, that's a scientific term, well, not quite, those very 24 hours to respond, 72 hours to report something to someone, and so on. And so we enter that phase where, on the one hand, we need to investigate, which we can do, but on the other, it's useful to help the security officer find the point to dig from, to narrow the circle a bit. Right. How does this usually happen? Usually, what can a security officer do? Most often, using one tool or another, he uses the information-system-based approach, as I call it. What does it mean? Here we have that same overloaded interface with ugly buttons, but deadly effective. So, basically the data-centric security approach: every little action of the user, when it's intercepted, is intercepted with a bunch of attributes, that is, the content part, what exactly he sent over some channel, and a heap of attributes: who was CC'd, how many attachments, what types of attachments, IP, MAC address and so on.
So what we're interested in, most often, in an investigation, is the account under which certain actions were performed by a certain user. And that's great, it works, it seems very convenient, and it can be applied in different guises, including, following the right advice, sometimes it really is better to draw a graph, because in a table the eye gets blurred. That is, some connections shown in a table are not at all obvious to the eye, and it turns out it was some mass mailing from a single external contact or something else. That these 10 rows are actually a bush, with 10 emails fanning out from a single node.
However, the main problem with the information-system-based approach is that we have to be somehow integrated with that information system, so that as fully as possible, ideally at the driver level from the OS, we can pull the user's actions. Before you is a piece of, no longer a DLP, but a DCAP system, Data-Centric Audit and Protection, where the actions of the user with a given file are recorded, that is, at the driver level. Why is this necessary? Because an agent, for example, can't always get everything for you. That is, the agent of one system or another, an agent, well, in my terms, a DLP system agent is very good on Windows, almost as good on various Unix-likes, not very good yet on macOS, and practically useless if we're dealing with cloud logic.
So if a user opened the already-mentioned Google Workspace, opened Docs in that Google and started writing a plan to take over and completely wreck the company. Right there, in the cloud text editor. From the agent's viewpoint and the actions we record for the user, what will we see? With a keylogger, say, we'll see that he's typing in some process or whatever, but a keylogger captures a stream. If the user got to item 10, then went back and fixed something in item 2, we still end up with gibberish, on the one hand. On the other hand, the main tool, if the user did all this in a browser, it seems we'd need to hook into HTTPS, the transport, intercept all the packets and see that the user was writing this bad document.
But we can't do that, in the sense of assembling the whole document, because in that cloud logic, only a small state change is sent in each packet at a time. So a person starts writing a plan to take over the company. The first packet carries some technical junk, which font, which whatever, and the meaningful part: well, "PL", two letters. The next packet carries "AN", another one, something else. In the end, formally you have the entire capture, but no complete document to analyze. And where does it exist? It exists if you're integrated via API with that same Google Workspace, with VK, with Yandex 360, with M365 or Office 365 via the Graph API, and so on. And obviously, what's the problem in this case?
Well, the problem is that you can't integrate with everyone out there. So besides the information-system-based approach, I consider two more approaches where steganography can help. The approach I've called user-session-based, that is, no matter which app the user is in, no matter how he does it, through the cloud, simply through VDI, and so on. What do we get in this case? You can't see it here, low contrast, but, for example, at the driver level certain watermarks are rendered, that is, on all monitors and so on. Where did such a solution come from in the world at all? From the problem of photographing the screen with a smartphone. That is, that's it, keeping any logs is useless. The user has access to this file, this information. He got access to it legitimately.
And at the same time the DLP system is powerless too. That is, access is handled by all sorts of things, including DCAP. DLP is responsible for controlling information in motion, if the user decides to forward these files, but the user doesn't forward anything, he opened it and took a photo. So, starting from this problem, we got the user-session-based approach. That is, as one option, rendering watermarks, that is, some technical information is displayed, and so on. What for? Precisely so that, if we move along the incident lifeline, that is, when a leak happens and some screenshots surface, those taken with Print Screen or shots taken on a smartphone, then a screenshot with a hidden watermark will speed up handling the incident. That same security officer will take it, see it, bring out those watermarks, that is, the username, the machine name, date, time, text or whatever we decided to display in that watermark. So this approach is used on the Russian market too.
Here's an example from, so to speak, colleagues in the trade. This is some public screenshot that I pulled. This is the company EveryTag, which does essentially the same thing. Here they say they can stuff a whole bunch of different attributes in through whitespace. That is, that very steganography, where each document is unique when a user requests it, thanks to certain shifts, either in line spacing, paragraph indents, left-right-up-down margins, and so on. And thanks to this, essentially, later, if a screenshot or photo of such a document surfaces, or a corresponding printout, you can extract the hash and see that it was John Doe who leaked it. So this is needed in addition to the information-system-based approach. However, there's an obvious problem here too. Fine, we've essentially solved one task, but there are more problems.
Thinking about these problems, I identified four of them. So the first problem: we have technical limitations by format. You can't get into every format, not every format can you stuff steganography into, essentially, apply steganography. Although maybe you can. Still, there's a problem with this. And what do you do, say, with audio files? That is, fine, text, that's the approach from the previous slide, text is clear, but audio, when it gets leaked, what do you do there? Unclear. The architecture of such solutions has problems too. Architecture in the sense that ideally you need to always provide each user with a unique copy, a unique version of the file at any moment in time, so that it all keeps changing. And again we get the technical limitation by format on top, plus the fact that it's something client-server and so on, unclear.
Technical problems with detecting the marks, the Achilles' heel. For example, our watermarks, which can simply be spread across all monitors at the driver level and so on. If you make them transparent enough that they, basically, are invisible to the user's eye, they show up perfectly if a screenshot is taken. However, if a smartphone was used in collaboration with the Russian hero Peresvet, that is, if you overexpose the shot, make it too bright, then naturally these marks disappear. You need to come up with something more robust. So there are also technical snags with detecting the marks. And there are also, in my view, exceptional situations. An example of one of them: when we say, embedded a mark, a file, written it in somehow, and a person opens this file and simply creates the same one beside it. Trivially, if we take a text document, a person legitimately opened the document, doesn't photograph it with a smartphone, doesn't copy-paste information out of it, but simply creates a new, completely unique document, astonishingly, letter for letter, with the same content.
What do you do? Nothing. And it turns out that here steganography isn't our helper either, in terms of how that mark would suddenly end up in there. So, basically, given the limitations presented, we arrive at this idea. This is an unsolved problem. I didn't come here to sell you some of our elephants along the lines of "look, you've been living wrong all your life, and now we'll live right". No. Nevertheless, the popularity of the steganography approach has grown in the last two years, demand has grown, the number of requests we get from clients has also grown, saying let's implement this too. And attempts to solve this problem lead to the following conclusions, in my view.
On the one hand, it's unlikely we'll manage to make something universal that would apply to all information transfer channels and to all formats, first and foremost, of information transfer. And a third variable is how to embed such a mark, or whatever it may be. I mean, well, steganography, what is it for us? It's when we've hidden one piece of information inside another, let's say, well, very simplified. As a result, for different channels, for a printer, say, it's still more convenient to use something of its own. For audio files or engineering drawings, better to invent something of their own. For text information, some third thing will work well. First conclusion. Second conclusion. Can we, in principle, still try to make something universal that would cover at least the majority of situations?
In my view, yes. From the angle, again, of interpreting the term steganography, I believe these could be labels. But these labels cover the cases where the information hasn't leaked yet. But at least we won't let the information leave the protected perimeter at all. By labels I mean labels that can be set on the file's address in the system, or labels that are written directly into the file's metadata. That is, again, we're coming back to the same watermarks that, for example, are used by the makers of software for creating various deepfakes. That is, they made a kind of global agreement and committed to developing their tools responsibly. And if anyone uses their tool to create audio files, in the spectrum or wherever, in short, their own watermarks get embedded.
Then detectors detect them well. So, accordingly, there's the same option with files: based on certain rules, either use an automatic label, or in parallel also use another label, this one, our manual labels, that is, a label that is written directly somewhere into the file. Naturally, hidden, naturally, so it can't be ripped out of there, naturally, so that it's inherited. And in this vein, steganography won't fully solve the leak problem for us, but combined with RMS logic it will, in principle, let us arrive at the following situation. That is, roughly speaking, so that a file, when transferred outside the protected perimeter, always goes in encrypted form. And if it reached the wrong recipient, well, that information couldn't be used then.
And the encryption would then be based on the label. That is, we essentially take the well-known attribute-based access model, attribute-based access control, ABAC, and bolt encryption onto it, onto that logic, essentially reinventing that Microsoft RMS which once existed, and it turns out there's demand for it again now. So those are the ideas, and in my view the Russian market is moving towards these ideas. We're moving here too, with varying success across different channels. And I'd like to hear the opinion or counterarguments of those present on this as well. So steganography isn't such a scary thing, in my view.
It will bring certain benefits, but you shouldn't think it's the one tool that will solve all problems. It's just a supplement. My personal view: in principle, the best incident is the one that never happened. So you shouldn't let it get to the point of execution at all, but catch all the insiders on takeoff, while still at the stage of forming intent or the information-gathering stage. That's all from me, thank you. —
— Alexey, thank you very much. Right, let's move on to questions. —
— Hello, thank you very much for the talk. First, this isn't even a question, more of a small addition regarding audio, because audio watermarks have existed almost since the 90s. And the second question. What's your view on the mechanisms used in Canarytokens? For example, that same zero pixel embedded in a document so you can later track whether the document was opened or not.
The zero pixel here, I understand it, but I take a negative view, and here's why. Essentially, we then create the need for an active action, that is, a risk of exposure. Roughly speaking, I even had an exercise like this in my lab, this spy pixel. A one-pixel image is embedded in a document, and if it's leaked and someone opens it, it phones home to some service like an IP logger. So what's the problem? The problem is that as soon as a document like that gets opened, something starts flying out, something starts calling home somewhere. And maybe the security officer doesn't want to give himself away. So it's a useful thing, you can use it, but you need to do it wisely, and that kind of functionality should be something you can switch off, I think. —
— Thank you very much. —
— Okay, colleagues, any more questions? I see no hands. Alexey, thanks a lot. Oh no, the speaker again. Alexey, thank you very much. Let's give the speaker another round of applause. That was great.
19. Oleg Bezik (Digital Research Laboratory) — “Automated government systems under the forensic computer expert's microscope: problems and solutions”
Scheduled 16:45–17:15.
Moderator's introduction
So, our next speaker is a real, genuine forensic expert. And you'd be amazed at what ends up under a forensic expert's microscope. Please welcome Oleg Bezik, who's going to look at a completely new topic. A round of applause. Oleg is a regular guest of ours, a regular speaker. —
— A resident, you could say. —
— Okay, this is forward... no, that's back. There we go, great.
Talk and Q&A
Greetings to everyone, ladies and gentlemen. I take it you're the survivors, right, who held out to the end of day two. Thank you very much. I hope I'll tell you something interesting and useful today.
I've been introduced, yes, I really am a forensic expert, I do forensic computer examinations, I do it at my own company, I'm the founder and CEO of a company called Digital Research Laboratory. Since 2016 we've been performing forensic computer examinations, and as of this year, 2025, we've also become an IT company: we're developing software for comparing source code in the course of forensic examinations.
A couple of words about me. I've been doing forensic examinations since 2014, I worked at Big Four firms, I hold several international certifications. Now I'm growing my own company. The company has existed since 2016. Over the last three or four years we've built up quite a lot of experience, including in performing forensic computer examinations of automated systems. Today I wanted to tell you about that, share a bit of our experience.
Let's start with what automated systems are in the first place. I hope everyone here uses Gosuslugi. That's one example of an automated system, one that gets developed and is used by a large number of people. There's also a category called state automated systems. What is that, essentially? It's a hardware and software system that automates some activity. State automated systems automate the activity of some specific government body, or some function. For example, GAS "Pravosudie" automates a function of the court system.
GAS "Legal Statistics" automates the work of prosecutors, and so on.
Why am I telling you about this at all? Because of how a state automated system gets developed. As a rule, it's a state contract, if it's done for the government, or else a contract between a customer and a contractor. The contractor is usually some legal entity that has a staff of IT people, developers, security people, who can write a program that meets the customer's technical requirements.
And very often relationships like that end badly. They develop the program, and it doesn't meet the spec requirements.
The value of such contracts is very high. Say, developing a system costs 150 million, 500 million, maybe even several billion rubles. And when the customer gets a system that doesn't work, while paying several billion rubles for it, well, that's not great, honestly, not much fun. And as a rule, situations like that always, well, most often end up in litigation, mostly it all happens in the commercial (arbitrazh) courts, legal entities suing each other, and sometimes it even gets to criminal cases. And in cases like that, 99% of the time, even 99.99%, a forensic computer examination gets appointed. Because judges understand nothing about how complex automated systems get developed.
And in order to figure out whether the program really fails to meet the spec requirements, they order a forensic examination. A forensic examination, I hope everyone here knows what that is, but just in case, let me remind you. Essentially a forensic examination is a way of bringing specialized knowledge into court proceedings, in our case IT knowledge, since it's a computer forensic examination.
On top of that, in this category of cases, where you need to check an automated system's compliance with the spec, they also bring in appraisers. What for? So that the appraiser can calculate how much the work is worth that was done correctly, or how much it would cost to finish the system, to bring it up to the condition stated in the spec. And then it turns into a complex, multi-discipline examination, that is, when the examination is done by experts in different specialties, with different areas of specialized knowledge. But we won't be talking about valuation examinations, that's a bit more probably for some other conference.
For now let's talk about computers. Here are a few examples of such cases, fairly serious cases. We worked on them, it's our experience. We did the examinations in these cases.
What is the purpose of such examinations, most of the time? First, to make sure that the program really does or does not meet the technical specification. Second, if it doesn't, to understand what specifically doesn't work and why it doesn't work. Because the reasons can always be different. It's not necessarily the developer's fault, that they're so incompetent. Especially when it comes to state systems, big complex systems, where the developer-customer relationship has to be practically like family, because the customer dictates the conditions, and the contractor has to react to those conditions very quickly. And sometimes there are simply communication problems, and because of that it all spills over into lawsuits. That happens too.
What other purpose? Again, the purpose is defined by the court's point of view. We think like a judge: what does the judge need in order to, say, rule on a claim.
So, the presence of critical defects; we'll talk more about what those are. And then cost, which I already mentioned: assessing the cost, how much was done correctly, how much needs to be reworked.
Roughly, these are the typical questions that get asked in this type of examination. They're actually more or less always the same, but it depends on the context. Of course, sometimes there are other questions, but mostly they look like this. That is, do the results comply with the requirements of the state contract, the agreement, the technical specification and the detailed specification. So, the contract has an annex, and as a rule that's the technical specification. And that specification describes in detail what the program is supposed to do. That is, what the automated system should do, what functions it performs, what results it should show, how it should show those results, and so on.
It's all described in great detail. But sometimes it's actually not that detailed. That's also one of the problems. We as experts get asked: does the program comply with the technical specification? We open the spec, and it's ten pages. Well, as they say, without a good spec, the result is anyone's guess. Well, that's roughly what this is about. But most often, when it comes to developing large automated systems, there, of course, the spec is very complex and well thought out, there are lots of items, the customer is very meticulous about drafting the spec and getting it approved. And as a rule, roughly speaking, here the customer is protecting itself. That is: we'd better ask for more, in more detail, than ask for too little.
But sometimes it's objectively different: the customer approves the spec and then makes demands that aren't in the specification. That happens too. And the developers often fulfill those demands. Well, to please the customer, it's big money, nobody wants to get into a fight. They fulfill them, but it still doesn't help, everyone still goes to court and gets examinations done. Right. So, next, the second category of questions is about cost. The cost of the work actually performed. For cost assessment there are certain methodologies, there's the Moscow DIT methodology for estimating the cost of software, there's the so-called COCOMO method, they're roughly similar, and they really do let you get more or less the same results, but that's more of a valuation matter.
I'm not an appraiser, so I won't go into it, I won't make things up for you. Right, so next is the category of questions about critical defects. This is also a very important point, because not all defects are equal. We often see this: there's some set of defects, and the customer of the system says, look, the contractor didn't finish this, and this, and this, and this. We start looking, and it turns out that this, this and this is about a week's worth of work, and basically the contractor is ready to do it, no problem at all. And sometimes it happens that the contractor says, no, everything works on our side, it's all fine, all good.
Yes, you sort of look through it, look through it, all the functions are implemented, but one function, the one and only function this software was created for, simply doesn't work. That's it. So everything looks nice, the buttons are blinking, like Nikita was telling us, everything's green and red, the filters can be configured, customized and so on. But the one report that matters most, the one that's used, say, for some important purposes, it doesn't get generated. Why? Nobody knows. And so what you get is a screw-up, a defect, that prevents the system from being used for its intended purpose. And it turns it, essentially, into a useless pretty toy that costs a lot of money but in the end leads to nothing good.
And it's made worse by the fact that people still have to work. System doesn't work, but people still have to work. Generate reports, deliver some result of their activity. And people end up having to do the very thing the program was created for, which is the whole reason it's built: to automate that activity. People are, in fact, doing it by hand. And people are also dying on that, dying from overwork. So critical defects are an important matter that is always checked in examinations like these. Well, almost always. And then there are the questions of whether a defect is remediable or not. That's also usually what the court is interested in, because it's one thing whether a defect can be fixed, and another thing when it seems fixable, but in order to fix it you need another contract just like this one, for another billion rubles.
And then it turns out the contract simply makes no sense anymore. So that's as far as remediable and irremediable go. That is, an irremediable defect, as opposed to a critical one, an irremediable defect is one that is simply impossible to fix, or not economically viable to fix.
Bless you. Next. And the last question: for what amount of the contract price was work of inadequate quality performed? Well, that's also more of a valuation question, really. Right.
Now I wanted to share the pain points we run into when doing examinations like these. Pain point number one is that methodologies for forensic examination of automated systems, you know, any kind of agreed, publicly available ones, simply don't exist, unfortunately. Well, we've already done a lot of these examinations, and for ourselves we've developed a methodology of certain steps that we take: roughly, we build a database of the objects submitted to us, check that all the deliverables are present, check that those deliverables meet the ToR requirements, and so on. So it's, let's say, just an algorithm of actions; for now, as a methodology, we haven't packaged it, but maybe someday we will.
It's further complicated by the fact that every system has something custom-made. It's like, you know, you can build lots of single-storey houses, but if someone wants to build themselves some unique palace, well, that's the story with automated systems. That is, there are no off-the-shelf automated systems sold in bulk, and for their requirements the customer always hires a developer to create something unique. By the way, there's another problem here: estimating the cost of all this work, because, say, one developer can do it for 100 thousand, another for a million, and a third for a billion.
And how to assess that is sometimes difficult. Problem number two: the huge volume of objects to examine. What are the deliverables in examinations like these? Documentation, technical documentation, meaning the ToR, the specific ToRs, the user manuals, all sorts of explanatory notes, and, in short, a whole million documents. By the way, the photo shows the object of examination from one of our cases. See, one box didn't even fit in the shot. That's just text documents on paper. A computer forensic examination. It's supposedly a computer forensic examination, but in the end it feels like it isn't a computer one at all.
That's one category of objects. What else is there? The program's source code. An automated system is, after all, programs. Programs are written by developers in some programming language. Source code is usually delivered on media, either on flash drives or on optical discs. Just blank discs that they burn the source code onto. Then they come to us for examination and we study them. And that source code is unbelievably huge. That's millions of lines of code.
What else is interesting here? The system itself, the automated system, as a rule, especially in government agencies, is designed for a huge number of people, for 5 thousand people, for 10 thousand people. I think the lighting just changed, or did I imagine it? Oh well. And accordingly, the functionality of these systems is complex. Imagine, we have an examination now, we did it, well, we've already finished, now we're writing the report. There were a thousand requirements, 1,041 requirements that had to be checked. So it's just an enormous amount of work, and specifically, when checking the functionality, how does it go: we sit down at a computer and, basically, open this automated system and start looking at what works and what doesn't.
We end up as testers of sorts, the role of testers. Not sure if there are any testers in the room. Anyway, we're similar in that respect. And what else? What other problems?
An examination can take a very long time. We have examinations that took us two years, a year, a year and a half. It's an unreal amount of work, and it's very difficult work. And, to be honest, it's expensive. Really expensive. Because, as a rule, examinations like these aren't done by one person. With us it's usually done by a panel of 3 or 4 people. We go, like to a job, to the office of the customer where the disputed information system is deployed, and sit there checking it. And the parties, the plaintiff and the defendant, come along with us. Our record, I think, was 9 people in the room. Besides the two forensic experts, seven more people, representatives of the defendants, sat with us, argued, quarrelled, told each other to go to hell, and we sat there watching all of it.
Quite a show, of course, but on the whole entertaining. You probably know that the parties may be present during a forensic examination. If you don't, here's a life hack for you: if anyone ever ends up in an examination, you can be present while it's being conducted.
Right. So those are the kinds of problems we run into when examining automated systems. And I wanted to give a couple of tips to those who do go into an examination, maybe someday it'll be useful to someone; I don't know how well this audience fits these tips, but if anything, pass them on to colleagues, lawyers or IT companies that develop automated systems, maybe it'll be useful to them someday. But on the whole, the main thing, of course, before entering an examination like this, is to understand in general what objects you currently have on hand. So if you're, say, the customer's representative, the contractor should have handed the deliverables over to you. Documentation is usually provided on paper and on media, on discs.
The program's source code and the deployed program itself. So those are the three main objects. You need to see if you have them or not, what state they're in, if... Oops, my... Ah, it works. Something flickered. Probably tired too.
And naturally, in such cases the objects of examination also include agreements. Agreements, contracts, there are a lot of curious clauses in there too, which we, as forensic experts, are also obliged to check.
The second point: it'd be a good idea to prepare the questions for the expert, yes, I gave a rough, approximate list yes, that you can rely on, but the devil is always in the details, and it all depends on the context, on the system, on the functionality that isn't working; that is, you can ask questions about individual modules, you can ask about the whole system, in short, there are many different options, but the basic concept for the questions, I've already given it. And third, of course, you need to choose the experts. The experts should still be... These systems are complex, automated, and as a rule the examination is long and difficult, so you need people who are professionals. I firmly believe that only professional forensic experts should do examinations.
I hope nobody takes offence if there are experts here who aren't forensic experts by education. And experience, of course, experience, because experience in such forensic research is very important, because once you've developed, well, a certain trained eye, yes, you can solve problems more precisely and more quickly. Right. Oops, where's the clicker? Here it is. That's it, colleagues, I'm done. Just one more thing I wanted to mention. We've made this checklist, a checklist-cum-memo on preparing for an expert report, for an examination of automated systems. You can go there and grab the PDF for yourselves.
Maybe it'll be useful to someone too. Right. That's all from me. Thanks a lot. Thank you, Oleg. Oh, so many hands.
You know, questions like these, about examining all sorts of systems, have been coming up for over 20 years now. When I worked at the Forensic Science Centre (EKC MVD), I tried to avoid this. What's the point? We had to examine an SAP system and banking systems, but never in our lives did we take on a question about testing. Let's look at it this way. There's economic forensic examination, accounting forensic examination. They fought to the death not to be made to do audits. That is, a full check of a warehouse, reconciling something else, that's not the expert's job. The expert's job is to check and confirm: doesn't add up here, doesn't add up there, why it doesn't, where it's recorded. And here you're describing it as if the system was built and nobody at all, there was no test programme and procedure, no acceptance of the work, and you start everything over from scratch. What is that? In principle, you should be checking specific questions that were identified: here's the discrepancy, right here.
A year ago we held, well, at Interpolitex we have our own event, where Mr Muzalevsky spoke, among others. We listened to six talks on the examination of automated systems. The first question that always comes up, and today Denis Aleksandrovich has left, he had his own problems to deal with, he talked about how the terms of reference should be prepared, what to pay attention to. And as a rule, the problem questions arise because the customer doesn't know how to formulate the tasks. Agreed. And then, you know, what simply astounds me. I can examine. Can you examine Microsoft Windows for ToR compliance?
That's too complex a system. So up to what system complexity do you go? Because, say, pilots are trained to fly on aircraft simulators. Can you assess the quality of that system without being a pilot? And you're unlikely to find the person who can help you with that from a professional standpoint. We're not specialists in, say, banking, so assessing how well that program has been written, is something you basically can't, or whether it'll work that way. Actually, as I say, one very correct, very good thing was said: when an examination like this is conducted, specialists from one side and the other should be present. They agree among themselves whether the actions we've taken are enough to check whether this program works or not. With SAP that's exactly how it was, because, well, a specialist who knows all of SAP just doesn't exist, and the consultants were configuring it, and actually this was the critical task: does the handheld terminal work?
I had a task like this too, where: we've been working for three years, you haven't finished it, and it doesn't work. And he goes: hold on, let's open it up, tick this box in this table, and over here, come on, check it: it works. And that's how it was. And so I don't understand when you answer qualitative questions. Quantitative ones, yes, but in a computer forensic examination, being responsible for the quality of all those actions, well, that's quite difficult. And then, look, they ordered a bicycle, and got a scooter. They built it, spent the money, but it's a scooter. —
— And so a valuation examination, damn it, what's it going to value? Sure, you can value the cost of the work, but nobody needs that. And in general, the market value of a public-services system, apart from the state itself, nobody needs it. Only the state can say, well, I'll pay that money, or I won't. Market value there is a really tricky thing. Unless you go, sort of, with the cost approach. There are many contentious issues here. Volokitin and I basically agreed that we'll work together on this and maybe come up with something. But for God's sake, don't take on all the testing. That's a road to nowhere.
Thank you very much for that comment. I wanted to respond. Can I, now? One second. Look, first of all, I wanted to say that, first, we conduct examinations on the questions the court puts to us. We don't define them. So the questions I listed are the questions the court puts to us. And the court needs to know whether the software meets the ToR requirements. What does the ToR say? The ToR says there must be such-and-such results. The results must comply with the detailed ToR. What does the detailed ToR say? We open it, and it says the program must do this. This isn't about SAP, this isn't about Word, this is about programs that are developed from scratch, custom-built for the client. That is, information systems of some kind.
For example, we did an examination of the State Automated System of Legal Statistics. It's a huge system that a large team developed over several years, from scratch. It's not off the shelf, and there are specific requirements for that system. And we can verify those requirements, as forensic experts, because they are set down in the detailed ToR.
And there's a test programme and procedure that we can use as a basis. —
— And how do you do that in two weeks? What's the question asked? It's whether it meets the contract requirements.
— The question needs changing, work with the parties.
— But we're not judges, we can't come to court and say, Your Honour, change the question.
— You can, why not? You can, and always could.
— I don't know, I've never come across that. Never.
— That's strange, really strange. All of that can be done.
— I just don't understand on what grounds we'd change the question. We come to court and say, court, you're thinking wrong?
— Yes, you file a motion saying that this isn't quite the right approach. Well, if you want to pay 5 million for an examination, go ahead, but you could pay 200 thousand, or 500 thousand.
And why would I want that? Ah, exactly, that's a different question. I don't get it, why would I want that? Then the approach is clear: don't deny yourself anything for the money paid. Of course, the court wants to know it all, and we'll tell it all. —
— Oleg, I wanted to say — my name is Barannikov, Sergey Nikolaevich, I'm also a forensic expert, and I've actually taken part in acceptance. So examinations like this arise when, under Federal Law 44-FZ, a state or municipal contract is concluded. And then the customer, to shift responsibility off themselves, so they don't go to prison for accepting God knows what, brings in computer forensic experts, among others, so that, step by step, the whole set of information documentation, the program code, with hashes, is checked for whether it complies or not. It's genuinely very painstaking work; I'll come show you later, I had mountains like that too. That was the state information system of Rosreestr. —
— And I really did sit there for several months, working on it, and then we reported that it complies, it complies. So such examinations are carried out, but not as part of the 44-FZ check, before it goes to court. This is what customers need to be told, so they cover themselves, because they really accept God knows what. The work is extremely interesting, it includes load testing and other tests, like that, to verify the data that would be... I've had to take part in that too.
— I understand you, it really is This work is very well paid, very complex and costly. —
— Congratulations on having taken part in work like that. It really is interesting work. Thanks a lot. Colleagues, any more questions? Yes, I see a hand, coming. —
— Andrey, please, don't knock me down again. —
— Oleg, thank you very much for the talk. I was looking at the questions that get put to the examination. Some of them concern cost issues, that is, valuation. So am I right that the examination format is usually multidisciplinary, that is, you bring in economists who deal with the valuation questions, the amounts, cost and so on? Or do you basically do it all yourselves within IT, with your special knowledge? Thank you. Thank you, Andrey. Well, as I said, such examinations are mostly multidisciplinary. That is, if there's a valuation question, you need an appraiser. Not an economist as such, but an appraiser.
Someone with an appraisal qualification who can apply valuation methods. I hope I've answered the question. —
— Another question. You were talking about valuing software development. Well, I don't know, I'll find out — people making software for 20 years, they can roughly estimate how many guys will be needed, well, development, how many designers, how much time, even consultants can. Honestly, I don't know what formulas you'd use to calculate, because developing some super-high-quality software solution that covers the whole ToR and about which no examination would ever say there are problems, right, if the customer for some reason decides to put the contractor in prison, then, well, I don't know how to calculate that. I mean, there's no formula; good software is — a huge contribution, especially if it's security, testing, deployment, well, that's many years, it's done over years. So, as I am, I just don't quite understand how you say you calculated it, right. Everything else I can see how to cost.
And the second question: for two solid years we... We're proving he stole the money, he must go to prison. And meanwhile, is anything being done so the system gets put into operation and can actually be used? Does the task get solved, or how does it go? That's also an excellent question. You're getting to the root of it. Look, on valuation, disclaimer first: I'm not an appraiser, so I can't give details. But I know we apply the Moscow IT Department (DIT) and COCOMO methodologies.
I suggest reading up, if you're interested. I just can't tell you the details, physically can't, I'm not an appraiser. But what's the idea? The idea is that these methodologies assume you can evaluate a program by some quantitative indicators. Number of lines of code, the number of, what's it called, not functionality, something like milestones or so. So, accordingly, those are the methodologies that are used. Apart from that, what other methodology did we use in our work? But mostly the classic method is like: it cost 100 million, 87 percent was done, so that's 87 million, right. Well, that's what we often come across, but we don't do it that way.
Well, we go by the number of requirements and how many are fulfilled and how many aren't.
— Just a second, but here, unfortunately,
— the weight can't be determined, — that's the whole point, so you're calculating the percentages all wrong.
— Why wrong? — Well, because every module has its own development complexity, and so, well, the contribution of that module.
— But in any case we somehow need to explain to the court how much is fulfilled, how much... — First, you're explaining it wrong. Second, I've repeatedly taken part, as the customer's representative, the functional customer, in commissioning and producing various R&D projects, including software systems and hardware-software systems. And I know perfectly well that the price announced for that work was, from the outset, lower than what companies would have spent that came in and started from scratch. Any company has certain prior developments. With prior developments in the field, you can do something for that price. If you've never been involved in it, you'll definitely, well, not so much lose the tender. You may even win it, but in the end won't do the work, because in a year, a year and a half, two years it's impossible. How do you account for prior developments, what the contractor must already have? They don't just go off for some price and do something.
It's always amazed me when people try to evaluate what they don't understand. Number of lines of code, quality of the program. We had it at one point: here we consider this programmer good, high quality, this one not so good, this one a bad programmer. How do you evaluate that? I consider myself an average programmer, and an expert tells me... I'd have written it better, so he's a low-quality programmer. And then I look: oh, how nicely written. That has always astounded me, really.
I'm sorry, one more question.
Fine, you rate yourself highly, you graduated from Bauman MVTU and so on, but there are two parties in the process, the customer and the contractor. The customer and the contractor came to an agreement at the preliminary stage, when the preliminary tests were done. Basically, 90% of the tasks are settled And they only have a dispute over the remaining 10% —
— But it's not always that way. So which of them will pay you those 5 million For re-confirming, right from the start, the 90% that nobody is disputing Colleagues, I think we're speaking different languages I'm talking about a forensic examination, one appointed by a court ruling. A request comes to us, we say it will cost such-and-such. The court pays us, not a party. That's one. Second, well, where the court gets the money from is not a question for me. We should probably hold the roast on day two. I don't know why every time legal questions come up, it nearly ends in a fight.
As if next time we'll just… It's just such a debatable, interesting topic, you see. Yes, basically, I think that's exactly how this legal business works. I have Valeria Mikhailovna, our lawyer too, and every time a question comes up, I just walk away, because I don't understand any of it. Thanks a lot, Oleg. Let's see Oleg off, a round of applause. Thank you, colleagues.
20. Mikhail Inkin (STC, Speech Technology Center) — “From audio and video data to evidence: AI analytics, biometrics, and spoofing and deepfake detection”
Scheduled 17:20–17:50.
Moderator's introduction
Could I just have the clicker back. And our next topic is the closing one for today, but it's more relevant than ever, because, I think, last year everyone and their dog rolled out artificial intelligence. Today, basically, that has brought a lot of useful things, a lot of interesting things, but also all kinds of challenges and threats. Which is what we'll hear about today from a representative of conference partner STC, Mikhail Inkin. Please, Mikhail, the floor is yours. Let's give him a round of applause. The bottom clicker.
Talk and Q&A
Good afternoon, everyone. Yes, hello. We've got technical problems here. Can we go to the first slide?
Let me flip through it myself, we've jumped somewhere here. Well, now you've seen all my slides at once, that's not bad either. A teaser. Yes, a little teaser. Somehow we started from the end, not the beginning.
Thanks. Good afternoon once again. My name is Mikhail, STC, head of a project group. And today I'll be glad to tell you how we use our biometric technology to help solve forensic tasks. A couple of words about STC. We've been on the market for over 35 years. We make products and solutions based on voice biometrics, face recognition technology and intelligent speech technologies. We've delivered over 5,000 projects worldwide, including the Russian Federation and many other countries. I'd note separately that our algorithms today are internationally recognised, they regularly take part in international scientific and technology competitions and prove their effectiveness on the international stage. Some of those competitions are listed here on the slide. NIST, the CHiME Challenge.
We have fairly broad expertise in working with large language models, including Sber's GigaChat model and also freely available open large language models. On top of that, we have expertise in fine-tuning large language models to a specific customer's request, to their practical, applied task. And today in my talk I wanted to focus on how audio data can enrich an expert's work, what additional insights this audio data can give experts. And I'll probably start with the question of voice biometrics in general. And briefly, it should be said that, overall, voice biometrics can be divided into several main methods. There are automatic recognition methods, where the work is done in automatic mode with minimal operator involvement.
Today we produce systems that are language-independent and text-independent. Whatever language the speaker speaks, our biometrics will recognise them. And, of course, there's the principle that the conclusions of any automatic system need expert confirmation. Of course, following that principle, we also offer our customers both automated speaker identification methods and fully manual methods. They're implemented in our products. Examples of the technologies we use in our automatic algorithms. And it must be said, our technologies work on different types of microphones and perform quite well in different acoustic conditions.
Just a couple of words about speech and the origins of biometrics in general. It should be noted that more than 70 parts of the human body are involved in producing the voice. Every person's body is unique. Because of that, the voice each person produces is unique, which makes it possible to tell speakers' recordings apart by automatic methods, and by manual methods. For example, for more than 30 years, audio recordings, forensic speaker examinations have been used as forensic evidence both in Russia and abroad. I'll probably skip the question of voice vectorisation, skip the question of how automatic models work. I'll say a bit more about manual approaches to speaker identification. This is mostly a forensic audience here. Clearly, a voice has different representations. The voice, an oscillogram, a spectrogram, a cepstrogram – these are all different representations of the same recording of a human voice. And it should be noted that a voice has identifying characteristics, for example, such as the formants, which are examined, for example, the fundamental frequency, which, when you work with them, let you reach unambiguous conclusions in the course of a forensic speaker examination.
We also can't ignore the question of how difficult this task is, because, as one of the biometric modalities that is easiest to obtain, a voice is easy to record: set down a recorder, a mic, grab a phone recording. At the same time, this speaker identification task is quite difficult and it's complicated by a large group of factors. These are speaker factors themselves, technological, say, different microphones have different frequency responses. Communicative ones, well, in one case I'm talking to my child, in another with my accomplice, in a third with a gang member, say, if I'm a criminal. And my voice will sound completely different in each of those communicative situations. Plus there's the group of technical factors: SNR, presence of noise, the speaker-to-microphone distance. All of this brings great variety into the audio data our users may have. And, of course, let's note that today our algorithms can overcome these difficulties quite successfully and perform identification even in difficult acoustic conditions.
Our products use the following methods, the main ones for identification. The auditory method, listening; acoustic; the phonetic method of analysing the fundamental frequency contour; the spectrographic and automatic methods. And today I want to talk about how different STC products solve forensic tasks. Probably, on some practical case, you can imagine that there is a set of audio data. It could have been obtained in different ways. Maybe it's lawful interception of phone calls, maybe it's microphone recordings, maybe it's a seized device and voice messages were pulled from it. There's some body of data, and we want to analyse it. It contains various unprocessed voices, some dialogues, some voice messages.
What do our solutions let you do with these data sets? First of all, let's say that we offer products for all the key stages of the forensic process. Today the main focus will be on data analysis with the AVIS product, and on obtaining evidence. That's the IKAR Lab 3 forensic suites, and also on the Nestor AI solution. We won't touch on data collection now, because that's a separate topic, and one could talk about it at length too. Let's start with the question of data analysis. These data sets can be analysed, for example, on the AVIS system. That's our search and analytics system, which in automatic mode, from raw, unprocessed audio data can extract a large amount of information. What can you do, for example? For example, you can automatically get information about the speaker's gender, the conversation language.
You can get the recognised text. Today we support 18 languages. Russian, naturally, the CIS languages: Ukrainian, Kazakh and some others. And also, for example, foreign ones, Arabic. You can run a biometric search, translate text from a foreign language into Russian, and once you've got the necessary data, you can go on to run analysis functions, for example, cluster analysis, statistical analysis, and do link analysis. On the following slides I'll show a few examples of how such a system works. Here is the application's working screen, and it shows the process of how a user of the system filters a large amount of data.
The system finds only those recordings that match the search criteria. For example, recordings that contain a particular voice, that contain particular keywords. Here they're highlighted on the oscillogram. There's a message transcript. You can run link analysis on all these recordings. For example, produce analytical graphs like these. And as one example of an analytical graph, it's possible to show a link between topics, what topics the voices of interest to us were talking about. Here one such recording is highlighted. It's definitely not her. His second one was a young junkie girl.
Definitely not her. Here's an example of that recording. Maybe we could turn the volume up a bit for the next examples. Here you can see that the keyword was triggered. And, accordingly, we need to go back.
Thank you. And this keyword that was found made it possible to assign this conversation to some specific topic of interest. Let's skip the next few pictures. Next, I want to talk about how large language models help improve search and analytics systems like these, the ones based on voice biometrics. The main capabilities of LLMs are probably already widely known. I won't dwell on this slide and will show a few practical cases of applying large language models in biometric search and analytics systems like this.
One approach is plugging in LLMs, large language models, to extract the information of interest from a body of data. That is, using a specific prompt, the user can analyze the data set that was found and get a brief summary of all the recordings found, highlighting only the recordings that are of the greatest interest to them. So the user reads a short summary instead of listening to every recording and reading the contents of each one, understands which of these recordings are of the greatest interest to them, and then works with each individual recording for further examination. Another use case, it's also shown on the screen here right now, is summarizing the text of a conversation. The process is as follows. Text is obtained from the audio, then the text is fed into the language model, and, via a prompt, the result comes back as a condensed textual representation.
I also want to show separately on this slide, for example, that our experiments with large language models show that large language models are quite good at extracting the information that is of interest to us. As one example, I suggest looking at item number 5 here. Code words and slang used. So, for example, large language models today are quite good at detecting recordings where some kind of concealed language is used, where certain objects or items are being disguised. And this gives new insights to the users of such systems.
Of course, I should note here that search and analytics systems solve search tasks, whereas evidentiary tasks, for those we have a different suite, the IKAR Lab expert suite, which is where forensic audio examination is performed. And I want to talk about what this suite can do in the context of this task. From the AVIS system, we can export the recordings of interest to IKAR Lab and run a detailed analysis of the recording. For example, do a preliminary quality assessment, decide whether this recording is usable in a forensic examination, and then carry out a full forensic audio examination, starting with noise reduction, extracting the textual content, separating the recording by speakers, and conducting an identification study, a full one, using various methods, including searching for traces of editing and also searching for non-situational changes in the signal.
I want to draw special attention, I'm going to play some audio now, could you please turn the volume up a little over there.
One of the recordings I found in the AVIS system sounds like this. For some reason it seemed interesting, we downloaded it and want to listen to it. I'll play this recording now, let's listen. Listen, I just got into an accident on the M4 highway, just outside Moscow. A Beemer slammed into me, the car's totaled, I urgently need money. Perhaps
those of you who work with audio were able to notice that there are some unnatural notes here. But as our practice shows, most listeners actually can't tell a synthesized voice from the voice of a real person. What we just heard is a synthesized voice. How does this happen? A sample of some person was taken, uploaded to a system, and then the user of that system typed in some arbitrary text, in our case clearly of a fraudulent nature, and the system voiced the text the scammer wanted in the target speaker's voice. A clear, classic situation: a voice message is sent to the parents, the parents panic, urgently look for money, send it, and so on.
This is quite a serious modern threat, a challenge to forensic audio systems: voice forgery, synthesis, re-recording, modification of voice messages. And I want to note that IKAR Lab is now equipped with full spoofing detection functionality. Spoofing is the term used to describe a synthesized voice, or to describe a modified voice. And I want to show the process. Here on the screen we see that the software can evaluate both an overall score for the whole recording, the probability of spoofing, and the probability that a forgery is present on short segments. That is, there's a sliding window, and every second you see the probability distribution on the lower graph, the probability that a fake voice is present in one segment of the recording or another. Because scammers today can forge not the whole recording but just part of it, some most important part.
IKAR Lab makes it possible to detect such forgeries. Moreover, the latest versions of this product can determine not just the presence of spoofing itself, but also the specific vendor. Here, I hope you can see, the slide says that with a probability of 99.6% this recording contains synthesis produced by the vendor ElevenLabs. Today we identify four main vendors. In the future this list will be expanded. I'd like to note separately that spoofing in general is quite a non-trivial task. Detecting a synthesized voice is quite an interesting and at the same time difficult task. It's worth noting synthesis algorithms are advancing by leaps and bounds, dozens or maybe even hundreds of new synthesis algorithms appear every year, so of course there's a kind of race going on here between synthesis detectors and the speech synthesis algorithms themselves. And often in a forensic examination the question posed is not the presence of spoofing, but the broader question of verifying the recording's authenticity, because traces of synthesis are often quite difficult to establish, given how it keeps improving and given that modern algorithms detect it very accurately and very well.
So IKAR Lab lets you take a broader approach to this task, including answering whether there are traces of editing using classical methods, for example phase analysis, background noise analysis, and so on. On these slides I wanted to show a very brief history of synthesis, to get a bit deeper into the subject. This actually isn't yesterday's invention. The first synthesized voice dates to the late 19th century. There was a mechanical device like this that could pronounce two speech-like words, "yes" or "no". There's even a link to the article there, you can look it up if you're interested. And here are some mass-market examples of synthesis. We'll probably listen to a couple of these examples now too.
If possible, turn it down a bit, just a little, they're just going to be rather unpleasant to the ear. Here, for example, is synthesis from 1979, the first mass-market device with built-in human speech synthesis. Let me play it.
Well, it's all clear here, even a child would easily tell this is synthesis. Here, for example, is the synthesis level of the 2000s, let's hear a fragment too.
Overall it's already more like the speech of a real person, but there are no emotions here, no breathing, just some speech-like sounds. They're also easy to tell apart from real human speech. I'll skip one here. And here's the current state of things – this example is taken from news sites, it's from last year, when in the presidential campaign in the US, a synthesized voice of then-president Biden was actively used for smear campaigning, let me play a fragment.
I probably won't play the whole recording, and besides, it's in English. But the most important thing I want to point out here is that every year speech sounds more and more natural. And if I go back to that fragment we listened to, the one with the BMW, we'll hear breathing, simulated breathing, of course, and pauses, and intonation. The systems have learned to copy all of this from a speaker's speech, and our anti-spoofing system can defend against and detect such forgeries. I'll add that IKAR Lab has functionality for manual examination, so the forensic expert can manually detect and document the features that they can then include in their expert report to confirm the conclusions of the automated system.
And, as I mentioned, there are modules for classical technical analysis. That's phase analysis, resampling analysis, background noise analysis, DC offset analysis. These methods live on and still help forensic experts find traces of authenticity violations.
Another interesting example. We got these recordings from open sources. And here we'll broaden a little the topic of voice biometrics into the topic of multimodal biometrics, and look at what our system can do when it comes to faces. So, this is our last talk, and there'll be some other activities afterwards. To get you engaged a little, I want to offer you a small puzzle. I'm going to play two videos at the same time now. One of them is a video of a real person, and the second is a synthesized video. That is, it's a fake face, not a real picture. And I'll ask you to answer the question: which video you think is real, the top or the bottom one. Then we'll just vote and see how the votes are split. So please watch carefully. It's Jessica Alba. My colleagues here are prompting me.
Let's see what it looks like.
No speech here, just the picture. I'll play it once more.
Let's watch it again.
So the first question is this. Please raise your hands, those who think the real video is the top one. Thank you. Now, hands up if you think the real video is the bottom one. Well, roughly even, roughly even, slightly more voting for the top video. Should I tell you the right answer now or later? Right away. Yes, the real video is below, yes, the one on top is synthesis. You… Well done, those who said "bottom". Thanks to those who took part and said "top". The top one is synthesis. And I want to show how our product can help a forensic expert answer that kind of question.
IKAR Lab now has a module for detecting deepfakes. It's frame-by-frame analysis of faces in video. On this clip I'll play back a video from the real interface, how it happens. That is, the expert can analyze video right in their own copy of the IKAR Lab software.
The system runs a frame-by-frame analysis. For each frame it outputs an estimate of the probability that it's a fake. And on top of that a mask is overlaid that shows which areas of the face most strongly activated the network's neurons. That doesn't mean those parts of the face are fake. It generally indicates that they for some reason draw the most attention from the neural network. For this video, we can see on the right in the interface that we got a fake probability of 97%.
And in addition, the software estimates, it says "in tolerance" there, the in-tolerance probability. That is, it also estimated the number of frames where there are excessive head turns or tilts. So the software also checks whether the face looks straight at the camera, or whether it's tilted, turned, and so on.
For an original video, I'll show you the opposite situation too. This is a video of a live person, with a genuine image of Ms. Jessica Alba. And here you can see that the estimated probability is very low. This lets us conclude that this is a video of a live person. And the probability estimate is also output for all the analyzed frames, and also for those frames that turned out to be within tolerance. So that's the kind of toolset we're now offering our customers as well, because deepfakes, spoofing and synthesis are a really serious challenge today that we face, and that's why our solutions are equipped with methods to counter such techniques.
I'll add that the IKAR Lab expert forensic system also has a manual assessment module for video synthesis signs, so that the expert, in their report, in their expert conclusions, can include their observations in a structured form.
And probably the last example I'd like to give here. This is our Nestor AI system. Once we've searched the data in the bulk set, found some recordings or media files of interest, checked them manually, first got some conjectures, then suspicions, detained the suspect, proceedings take place. And those proceedings can be recorded, but they'll be recorded on camera, on a microphone. We also offer a system, the Nestor AI system, which can automatically produce minutes of official sessions. That is, it can work both with streaming audio and with files obtained from a system, say, a conference-call system such as Zoom, or dictaphone recordings.
The output of such a system is a strictly formatted record containing the transcript of the meeting, split by speaker. And I'll also note that we build large language models into these systems, which also make it possible to summarize the results of meetings, to produce short, condensed abstracts of these meetings, for easier work with such records. Here you see an example of the system's interface. So the number of speakers was determined, and the gender of those speakers. Here's the transcript, and now the output of the AI agent will be shown.
Here's the meeting record. And with prompting, of course, you can tune the output of these agents depending on the needs of our customers. As a conclusion, I'd probably say that we work at different stages of the forensic process and bring a new modality to the data you already have, allowing you to find new insights. Thank you. We probably have a little time for questions. Yes, of course we do. Mikhail, thank you very much. So, colleagues. —
— I have two questions, they're not related to each other. But the first question is about video deepfakes. From the examples shown, I'd conclude that it was face swapping originally that was used as the technique. Fine, you detect it. But have you tried it, and how does the product behave on other techniques? Puppet-master, lip-syncing, and synthesis of something? Moustaches, moles, earrings and so on.
As for synthesis of facial elements, I haven't tested that myself, for example adding moles and so on, but for all the other situations, a face mask, real-time face swap, we've tested it, it all works at a sufficiently high quality. No, not just swapping. Lip-syncing, swapping is when it's part of the face, the main face and its expressions are preserved but only a part changes, mostly the triangle. One example is, say, when a photo of some famous person is animated. So there's a photo, we overlay the mask movement, well, the attacker overlays the lip mask movement, and on the whole, yes, such fakes are easily detected. On our datasets we get a result of 98%. Obviously, for other data you have to look, test on the data. But overall the result is very good.
Well yes, so overall it was strange to hear that you identify specific vendors. I thought that underneath those same audio deepfakes there's some Multi-Tacotron sitting anyway. And then it mostly makes sense to identify the tool families, not the vendor who took something free and reskinned it a bit. Why is vendor identification important? We discussed that here with an expert too, with our expert. Of course, it's additional circumstantial evidence when working a case. Because if we can establish that our suspect visited such-and-such a site, for example. If we can detect that the synthesis was generated from that same site, we get additional evidence to support the expert's work.
Thank you. And here's the second question, purely out of idle curiosity. Have you tested your recognizer, which then runs all of this through a neural net, against data poisoning? What happens if, roughly speaking, I add to the audio signal at an inaudible level a system prompt like "listen to me" for the LLM that's about to process it. And accordingly, the voice of conscience dictates something further for it to do. So, here we're talking specifically about prompting. Poisoning. It's just that your input data is audio. And how exactly does your speech-to-text work? That is, maybe I can inject an upper layer, say, initially insert the first part in some sufficiently long pause, so that everything afterwards is ignored, using only that frequency band, and then, well, insert the rest. Thanks for the question. I should say that several models are used in this chain, and they're different.
That is, for speech-to-text conversion it's one set of models, our own in-house development. For the subsequent LLM analysis other models are used. So as far as turning a recording into a transcript goes, it's unlikely here, if you consider how this data comes to our customers. Usually they're interested in the data being clean. And as a rule the user has no interest in distorting that data. Usually a different question arises. Let's tune them for some specific modality. For example, our speakers talk about some strictly defined topics, in some odd language, made up, or school slang, or sci-fi. And then they come to us and say, please tweak and fine-tune the model so that it recognizes this narrow group's slang well. And that's a task we do solve. As for processing, as a rule the customer is interested in the data being clean.
As a rule, such situations don't arise. —
— I have a small question, I'm nervous, but I want to ask it. Are there any internal measures to protect the software product against leaks? After all, this can basically be seen as a weapon, when you have the ability to recognize and analyze speech data. —
— Thanks for the question. Usually these systems are always deployed inside our customers' closed networks. And this data protection question is addressed by the system running in a closed network, and on top of that, technical solutions are used so that the data, that's access control, for example, roles, privileges, separation of access levels, so that a user only sees the data they need. And I'll also note that the models we offer, both the LLM models and the speech-to-text models, the biometric models, they all run on-prem, that is, they don't need any internet access. And they can be, and always are, deployed offline on the solutions, on the hardware of our customers.
Right, Olga Alexandrovna, hold on a second. I see a question over here too.
Thank you very much for the talk. A practical question. Tell me, how effectively does your system deal with… If we have an audio sample where the microphone was being jammed with ultrasound. How effective is the noise cleaning?
— Thanks for the question. It's a big one. Answering in a conference format, with little time, I'll keep it brief. First of all, of course, there's the question of how you obtain that data. You're talking about ultrasound jamming or something. That's a separate topic in itself, one worth debating, because in our practice we very rarely run into situations where such jammers can actually cause any serious interference to good microphones. Let's start with this: as a rule, it's something of a myth that such solutions seriously protect speakers from being recorded. The second point to make: the algorithms are, of course, fairly robust to changes in the acoustic environment: SNR, reverberation, signal-to-noise ratio, reverberation level, and so on. Naturally, when developing these algorithms we always assume that the acoustic conditions the recording is made in, as I wrote there, may be not only a close-talk but also a far-field mic.
They may be far from ideal. So up to a certain limit, a certain SNR value, the probability is very high. Of course, to be honest, the lower the SNR, the more noise there is, the lower the probability of correct identification by an automatic method. But here we can say that there are manual systems, expert ones, such as IKAR Lab, where, using manual methods and other approaches, for example linguistic ones or something else, such problems can be successfully overcome. Thank you. And the second question is more commercial, I suppose. Your product Nestor, does it come as a device or as a license renewal? It's a hardware and software system. Different delivery options: microphones plus software, or just the software. So the software can use the microphones the customer already has installed, microphone arrays. We offer such solutions too. So it can be purchased as a software product. —
— Right, Olga Svetlanovna, I recall a question. How fast is your speech transcription? It's just that I know of systems that need a great deal of computing power to transcribe an audio file with good quality. Thank you. The question is clear, but again there's no simple answer.
I should say the following: transcription speed depends on the data volume, obviously, and the hardware. And, of course, mainly on the hardware. Here's an example: on modern GPUs, for instance, on server solutions, the speed-up of speech-to-text conversion can reach a thousand or even tens of thousands of times. Roughly speaking, a thousand seconds of speech become text in one second. So you need to sit down and look at the data volume and what hardware the customer has. So, in other words, transcription in real time… Considerably faster than real time. Excellent. And at the very beginning you said you support several languages. As far as I know forensic phonoscopy methods, what's in demand here is Russian, basically, wherever native speakers are. A phonoscopist can't do a forensic exam if they're not a native speaker of that language. So who needs analysis in other languages? —
— Yes, I'd answer like this: these systems… First, IKAR Lab has the option of manual transcription, so that a native speaker of, say, some regional language, if it's not supported in semi-automatic mode, could do the transcription manually. And second, IKAR Lab is exported to a number of countries, and those languages are quite relevant there. —
— Mikhail, thank you very much. That was the final talk and the final question of today's conference. Thank you.
Conference closing
Olga Gutman (MKO Systems). Scheduled 17:55–18:00.
Moderator
But overall, let's do this: the way we opened is probably how we should close. I invite to the stage the ever-wonderful Olga Gutman. Olga Vasilyevna, please, the floor is yours.
Closing — Olga Gutman
Well, these two days have flown by. They flew by very quickly. For me personally it was incredibly interesting to listen to all the speakers, and talk with you throughout our event, the whole conference. I thank you for finding the time both to attend our event in person and to watch it online. This time we set some truly incredible records for live-stream viewership. All the detailed figures will be available later in our post-event release. Thank you for your work, for the labour you put in every day. I wish you great, great luck and success in it. We, as always, will do our best to help you in your difficult work. I thank the speakers, a huge thank you for the most interesting talks, useful, the most important talks.
I thank our partners for co-organizing our event. And, of course, our team, thanks to whom this event took place. Let's give them a round of applause. Guys, please come up on stage.
[applause]
When you see us all separately at different spots around the event, it seems like there are very few of us, but actually, please take a look at how big the team is that's responsible for making everything that happened over these two days happen. And our photographer, who's standing there modestly, you can't see her. Thank you so much, guys.
[applause]
One second, maybe... well, all right, later.
While the guys are gathering for a group photo: at the end of our events we always thank the team that organizes them, but I've noticed a big omission, I think, on our part. Our development teams, who create our product, are watching us. Let's give them a hand, because, guys, thank you very much. You didn't see the roast yesterday, but believe me, so many kind words of gratitude were said for creating our product, for its constant improvement, for the support. We bow deeply to you. Once again, applause for our development team.
And for our guests, that's right. With that, the Moscow Forensics Day 2025 conference is declared closed.
But we're not saying goodbye for the rest of this half-year. Please come see us in Saint Petersburg: those who couldn't make it here from Saint Petersburg and the nearby regions can meet us in person on October 7. On September 24 you can see us in Astana at a partner event, and on October 2 we have a trip to Novosibirsk for the Legal Week, where you can also come talk to us, and Lera here is actively prompting me. Get ready in advance: a year from now we have our anniversary conference, and we're already preparing for it, lots of cool features, be sure to come. Once again, thank you very much for your attention over these two days. It was awesome! Goodbye!
[music]
[applause]