1
0 Comments

Malparse DevLog #2 – PE Parsing and Learning Design

This devlog was originally posted on my blog with a TL;DR Tweet Thread posted here. so check those out if you don't feel like reading the whole thing or want to see it on my blog!
--
The project lives!

I’ve been hard at work developing Malparse and all I have to show of it is some lackluster Figma designs and a broken back end. I’m dividing the devlogs into sections for now to show what I’ve worked on thus far in the many facets of bootstrapping a SaaS app together. I’ll start with Marketing and Design, then Back End,and Front End development and, finally, Data-Focused development.

Marketing

Since I’m trying to start this up on the cheap, one of the biggest areas I can cut costs, and conversely one of the worst places for me to cut costs, is marketing. It’s a hard balance I’m going to struggle to find, mainly because I absolutely despise marketing and I have zero successful experience in doing it. Even my web scraping course hasn’t made a dime off of the couple of short-term Google AdSense campaigns I’ve run for it.

What I am decent at is social media and p2p networking. I don’t like throwing flashy ads at people, mostly because I hate them myself, but I’ve been putting the idea of Malparse out to some people in private groups, and plan on using these DevLogs, my YouTube channel and Twitter to do more organic marketing. My thinking is, I’m trying to make a tool that really does some cool stuff. I’ll do some of the normal marketing stuff, but I really want the product’s capabilities and features to speak for itself.

So, word-of-mouth and #buildinpublic will be my primary marketing points, at least until I build up some revenue to start spending on marketing.

Design

Ugh… I hate design! It’s another thing I’m hopelessly bad at, but luckily there are tools like Figma that allow me to create crimes against humanity like this…

v0.1 Design
My horrible v0.1 design sketchup with Figma

Yes, I know it’s horrible. There’s a lot of work to be done here, alright?!

Figma is awesome, I don’t want my lack of design skills to make the tool look bad. Basically it’s going to allow me to design the site without writing a bit of CSS code, which is absolutely ideal since CSS makes me want to stick my head in a blender. Then, when I finally have to put code-to-screen, it generates most of the CSS for me, leaving my job to mainly be relegated to creating the right kind of HTML layout.

Thus far, I’ve just been learning the ropes of design and Figma both. Not too much else I can do but learn!

Back End Development

So, the current system model is this:

Front-End (React) -> Public-facing API (Express) -> Data API (Flask) -> Data processing (Python) -> Data Store (MongoDB) and File Store (Python)

It’s convoluted as hell, but I knew there would have to be some kind of translation layer because I want all the actual data processing to happen in Python. So for the back end, the Express Public API has to translate the user request to be readable by the Flask Private API which then does all the processing and all that fun stuff.

Thus far, I’ve got an initial DB schema for PE files laid out in Mongo. Of the CRUD operations, I currently only have an endpoint for data creation. I basically just wanted that up so that I could start working on the data portion, both because that’s the meat and potatoes of Malparse and because it’s way more fun.

Front End Development
I have intentionally avoided the front end because this is meant to be a data first app and I want a solid proof-of-value before I actually write most of the front-end.

Dead Simple Front End
The Dead Simple Front End

So, this is it. It’s a screen where I can upload files to be processed by the public API and Data API. When those two API’s are working (which is fairly rare…) it will show some initial data about the file you uploaded, but nothing big.

Data-Focused Development

Now for the actual fun stuff!

I’ve built out a parser for just about every field in the PE file spec to use later. The processing time is fairly slow, though, so I imagine I may have to get rid of some fields and figure out a better way to process the files. Since right now I’m focusing mainly on executables, one of the biggest ways I can save time is to minimize the amount of times I have to loop over the bytes in the executable. This is probably the most time-intensive operation since most of the rest of the parsing is just array slicing, which might get memory intensive but it’s not as time intensive.

The only thing I really have left is parsing and calculating the flags in a couple of fields. I haven’t worked with flags before so it’s taking some research to understand what kinds of data transformations I have to do before the bitwise operations, and then what bitwise operations I need to do to properly process the flag.

What’s Next?

After I’m finished with the flag processing, I’m going to work a bit on the front end to show the fields and make sure that they reflect correctly. Then I’m going to work on cache checking, so that if a file and its information is already in the database, I don’t re-upload and re-process it. I’ll probably have cache-checking toggleable for testing purposes, or I’ll have a toggleable “re-analyze” button on the front end.

After that, it’s on to parsing the ELF files. I knew very little about PE files before writing the parser, but I know even less about the ELF file format. That will definitely be a learning experience.

on March 1, 2022