This was originally posted on my blog
I've been in information security for over 5 years now. It's been a hell of a time, and I've been in various development and non-development role in some of infosec's biggest companies.
Over that time period, I've learned some phenomenally important lessons. None of them have been bigger than this:
Data matters.
I know, it's a stupid, two word sentence to be such an important lesson, but it's true! So much of my jobs over the last several years have involved building, maintaining and supplying pipelines of data. This data needs to be restructured and enriched from that source using these 15 API's, that malware is present on these three platforms that all contain different metadata.
It can be a pain in the ass!
So, I figured, why not create my own SaaS platform to further that pain in the ass? 🤣
Only kidding, obviously. The industry defacto standard tool that just about everyone in information security has to know and use is VIrusTotal. For those of you who don't know what VT is, it's basically Google for malware.
Okay, now that I think about it, that might not be super clear. 😅
Basically, VirusTotal is a place where you can upload malware (or what you think is malware) and have a super-smart system analyze that malware. It pulls out indicators of compromise (IOCs) like IP addresses, passwords, etc. and presents them to the user, runs the malware through a bunch of antivirus vendors to see what they think it is, and a bunch of cool stuff. It also allows you to search for file hashes (SHA256, MD5, etc.) to find malware that you may see on your system, create alerts for different strings and IOCs found within malware, etc.
In a couple words, VirusTotal is awesome.
There are a couple things that VT doesn't do, however. VT is built for all kinds of malware and files, and it does all kinds of analysis, but that, in itself, is kind of a weakness. It has to work for all viruses and files and have functionality that fits all of it. That means that certain niche needs that a malware analyst might want aren't present.
Here's an example. One of the most common pieces of malware that infosec has to deal with is CobaltStrike. CS is a modular, configurable piece of malware that's actually a legitimate piece of software sold to "white hat" hackers, or hackers that hack legally on behalf of companies or contractors that hire them. It's also frequently used by cyber criminals and hackers working on behalf of states.
One of the most frequently sought after pieces of information for CS is it's configuration file which can be extracted from each sample. This configuration file contains information that's often unique to each actor, or at least unique across different campaigns. VT doesn't extract these configurations by default, so information security analysts often have to do it themselves.
Pain point
Another common requirement is examining macros embedded in Microsoft Word or Excel documents. This is often a manual process and one that VT also doesn't do.
Pain point
Some malware has its own proprietary configuration schema where it stores important information like PDB's, command and control server information, etc. VT naturally can't do custom configuration extraction for every malware on the market (and, spoiler alert, Malparse won't either) so, we've found another...
Pain point
The final pain point is that VT is relatively limited by what values you can search for. You can search by hash, filenames, etc. but not necessarily values in the malware's configuration schema, for example.
Malparse is going to solve a lot of the shortcomings of VirusTotal, but more than anything it's going to be a big data analysis platform for the malware analyst. It's going to focus on a lot of the more low-level minutae that VT and other tools haven't solved at scale. More than anything it's going to allow for as many of those fields as possible to be searchable and indexable by the analyst.
Things like macro extraction and common configuration extraction will be available by default from the beginning, but I'll also be focusing a lot of the effort on building a modular, malleable method to develop new extraction techniques, at first by me but later by the individual analyst. That's going to be the really hard work later on, in really building this into a specialized tool by and for the analyst.
Let's not get ahead of ourselves, though... We need something to get us started, something to get some use and purpose, something to gather malware samples and user stories and bug reports and... lots of other things. That means the most loved, most often misunderstood and often, paradoxically, most hated term in SaaS startup history.
So on to the MVP. I've got it broken down into big features and small steps to get there. I've already got the beginnings of an API and since API and backend development are my strong suit, compared to front end development and design. So why not make it fun, at least in the beginning?
✅ File upload - File upload is done, or at least the beginnings of it. I have a simple button interface to upload a file from the front-end and send it to a server.
❓️ File upload -> File store API - Now that files can be uploaded, they need to be sent to another API to store and save the file and store its initial metadata in a database.
❓️ File analysis framework - Okay cool, so the file has been uploaded and stored and we've got some very simple metadata like file hashes, upload date and uploading user. We're not just trying to give the user file hashes, though, right? Now we've gotta start doing some analysis. The backend will queue up the file for analysis. The actual analysis framework I'm going to build in Python, because that's the language I do most of my security work with and, since a lot of this work is going to be fairly low-level, I think it's the most suitable language. So the user will upload a file, the file will be sent to the API, which will extract some minimum metadata and store the file on a filestore server before inserting the file into a queue. The Python analysis framework will service the queue, store information in the database mapped to the metadata that was already stored and notify the API that the analysis was finished. Then the API will fetch that data and return it to the user.
So that's the basic framework for the next couple phases of development, and that's where the pieces of the MVP fall in... or at least parts of it.
I'm gunning for a simple UI/UX at first. I want a pretty landing page with all of the marketing hooplah and user stories and all that. A user acquisition flow that will take the lowly curious browser off the net and convert them to a paying customer. A login flow that will take the recent converts and send them to a dashboard. And, finally, the dashboard, where all the magic happens. File uploads and downloads, analysis and searching and all of the rest of the fun.
I'm going to be very honest, I don't plan on this being particularly pretty. Design is definitely not my strong suit, and neither is front-end development. If this thing could entirely be a CLI SaaS service, I'd honestly consider it. But, instead, I choose an ugly, yet functional, UI/UX.
I'll be using React, both because I think it's the most based of the front-end development frameworks, and because I have the most experience with it. It will also decrease the amount of context switching I'll do between the front end and the back end I'm writing in Express.
Speak of the devil...
The back end, at least the API that routes requests and responses to and from the user, will be written in Express with database storage in Mongo. The queue will be maintained in Mongo, as will all the data stored about individual files and likely users as well. The data science bit, the actual examination of files, will be written in Python.
So, what features will be in the MVP?
Macro extraction - I want to be able to automatically extract macros from .doc and .xlsx files, store those macros as their own objects and store metadata about them so that they'll be linked to the original document file but also be stored as their own thing. This will allow an analyst to search for macros that are present in multiple different documents.
CobaltStrike Identifier - I want to be able to answer the supposedly simple question of "is this, or is this related to, CobaltStrike?" I have a feeling that this isn't going to be an easy task, but it's one of the many questions that I'll want to answer to facilitate the next step...
CobaltStrike Config Extractor - I want to be able to extract CS configs and store them as their own objects, similar to the macro extraction feature. This will allow you to track CS samples over time, search on different fields within the CS configuration, etc.
Metadata Extractor - This one is pretty self-explanatory. I want to pull out general metadata, such as file types, embedded strings, imported functions and DLL's, etc. from each file. This is present in VT, and it's super useful, but I want to maximize the amount of fields that are searchable. I want the user to be able to narrow a search by file type and imported functions so that they can look for files that have string "foo" and import function "bar()" so that they can cluster similar malware together.
I'll be following the #buildinpublic movement with this project. I'm doing it for a thousand different reasons, from marketing to fielding feature requests, but mainly I think it's a cool collaborative idea and helps keep my thoughts straight. So, if you want to follow the journey, you can follow me at any of the links below: