
Statfolio
Portfolio tracker and analytics for stocks and ETFs
Until today, Statfolio only encouraged orders newer than 2017. We could have theoretically supported trades as old as 2015, but with a bunch of gotchas.
We worked hard to extend this limit to 2010 and to do so we focused on two main areas:
- Sampling metrics before 2018 (TODAY - 3Y) on a weekly basis and daily after 2018. To make this happen, a lot of improvements were done under the hood to properly align the data.
- We added more than 50 million datapoints, effectively doubling our data store. Again, lots of improvements were needed to make sure our systems are able to handle to huge amount of data.
In order to be able to handle more data, we improved our indexing and data locality. One of the main effects is that tickers can now be retrieved more than 10 times faster, even if they are not yet cached.
This reduced the latency of some of the website pages such as /ticker from a couple of seconds to just a few hundred milliseconds.
It also sped up the processing queue and effectively halved the runtime. This way, we are now better prepared for another round of growth.
1 Like
Comment
Up until this point, Statfolio only tracked orders for USD-denominated stocks from NASDAQ, NYSE, BATS and some OTC stocks.
But since we now have customers from all over the world and this was one of the most requested features, we added support for the following exchanges:
- TSX and TSXV (Canada)
- LSE (United Kingdom)
- XETRA and FRA (Germany)
- PAR (France)
- ASX (Australia)
- NSE (India)
In order to do so, we also introduced the concept of currency exchange. So now, non-USD stocks and USD stocks are correctly aggregated into USD-based metrics. We can also leverage this to introduce another type of profit: "currency gain", besides the existing "dividend gain" and "capital gain".
1 Like
Comment
This includes both work from today and from 2 weeks ago. Our workers are processing a lot of data in background. Some of the tasks they perform are updateOrder and updateHolding, which as their name tells, are updating the data for orders and holdings. Such an update can happen either when a user changes (or adds/removes) an order or when a new daily datapoint comes for the price of the stock (one each day).
Before the performance upgrades, both updateOrder and updateHolding had runtimes between 8-15 seconds. That was mostly caused by the fact that getting the ticker data involves selecting a couple of thousands of datapoints (one point per day for the last few years) from MongoDB. We have over 22M datapoints, so finding the right ones is not a cheap operation.
In order to improve the performance, the following changes were made:
- Pre-caching the entire ticker data for each symbol after the first retrieval. So instead of selecting 5000 records, we're only selecting 1, which contains the entire data series.
- Today, I wanted to improve the runtime even more. I initially hoped for an improvement of 5-10%. I did the following 3 changes:
- used mongodb aggregations and projections to only select the range/datapoints I'm interested in, instead of selecting the entire ticker history.
- added an index over the symbol and exchange keys, so that the selection operation is faster
- started grouping the order and holding updates for the same symbols. So orders for MSFT are updated one after another, then orders for GOOGL etc. Doing so takes advantage of the internal mongodb memory caching from frequently read values (essentially avoiding expensive HDD reads)
After the improvement no 1) we reduced the runtime from 8-15 seconds to 2-4 seconds which was a huge win 2 weeks ago.
But with the today's improvement, we reduced the runtime from 2-4 seconds to 30-80 ms. That's way above any of my expectations.
So to summarize, after both updates, the runtime was reduced from 8000-15000ms to 30-80ms, a staggering 99.7% reduction. Let's just say that MongoDB never ceases to impress me.
3 Likes
Comment
At the time of writing this milestone, we are handling over 20 million datapoints for the stock prices alone.
We are also handling around 50 million datapoints for all the individual charts we offer to our customers.
Obviously, handling such amounts of data in real-time is nearly impossible. So to do it as fast as possible, we created an asynchronous task queue + a worker fleet that processes those tasks in the background.
Every time a user changes or adds a new order, the system does the following chain of operations.
- It updates the order meta data
- It updates the corresponding holding by aggregating all the orders for the same symbol in that portfolio
- It updates the portfolio to reflect the change in the holding metrics.
- It invalidates all the caches for that portfolio.
This exact processing chain is also used when updating the prices daily. Whenever a new daily datapoint is added for a stock symbol, all the corresponding orders are updated, and the above chain is re-run so that the portfolios get the new datapoint.
The same happens for stock splits, company mergers and acquisitions, reverse splits etc.
This pre-processing work, although intensive computationally, acts as a form of pre-caching for the entire website. So when the user wants to see the metrics under their portfolios, they are able to see them instantly, even though each metric has thousands of datapoints. Computing these in real-time would take seconds or even minutes, so its not feasible at all.
The largest portfolio we handle at this moment has over 1500 orders in it, and the pages are still loading in under 500ms for it. None of our competitors can offer such metrics and in such a fast way, so this is a huge win for us.
3 Likes
Comment
During the past few weeks we have been working to expand Statfolio to non-US markets, by including stock exchanges such as LSE from UK or TSX from Canada.
There are some difficult aspects to cover before we can safely move ahead:
- How to reliably do currency conversion. We don't want to just use the last exchange rate, but instead use the historic rate for each daily datapoint. The result is way better in terms of accuracy, but more intensive in terms of computing power.
- How to source the data? Getting financial data for US companies is easier than for foreign ones.
- Everything to this point was indexed based on the symbol of the security. With the international expansion, we will have to deal with companies that are listed on multiple exchanges with the same symbol, such a Fortis which is listed as FTS on both NYSE (US) and TSX (Canada). So we have to switch to (exchange + symbol) indexing. This is easier said than done, because at this point, the quantity of data we're processing is immense.
3 Likes
Comment
We started the website on the statfolio.net domain because at that time, statfolio.com was not available.
We now managed to buy statfolio.com, which we currently only use to redirect its visitors to statfolio.net. We don't have any plans for it yet, but we are definitely going to take advantage of owning this more prestigious domain in the long-run.
Moving the entire website to the new address is probably not a good idea in terms of SEO. So what I'm thinking is to use the new domain to offer some related products, such a financial APIs. We are currently sitting on large quantities of valuable financial data, such as stock prices, news, financial reports, company logos etc. All this data was manually checked for accuracy, so it can prove to be very valuable for other startups. But no concrete plans for the short term.
2 Likes
Comment
With the lifetime subscription campaign, we got featured on some of the largest media publications in the world, including IGN and Yahoo Finance.
This in turn drove a significant increase in traffic over the last couple of days. We were also able to sell more licenses after receiving all this media attention.
https://www.ign.com/articles/track-your-stock-portfolio-life-statfolio
https://finance.yahoo.com/news/stock-market-easier-portfolio-tracker-173000534.html
2 Likes
Comment
We partnered with StackCommerce to sell lifetime subscription licenses through their online marketplace + their partner network.
In order to support this, we introduced the idea of licenses. We also use these as awards to the community members that are contributing by reporting bugs and vulnerabilities.
We got our first $500 from subscriptions after we managed to sell a both basic and pro subscriptions.
We also started to get some positive feedback from customers. It seems that paying customers are way more involved in the product lifecycle. Glad to be able to serve them.
1 Like
Comment
About
As a stock investor I used to track all my investments in an excel file. Over time, as I continued to evolve the file I started to hit its limits. That's when I decided to implement the same functionality as a website


Comment