If It Has a CSV Export, It Has a Pipeline
My charger now does its own bookkeeping
My company car charges in my driveway, on a charger I own. So every month my company owes me money for electricity, and every month somebody has to prove how much. Log in to a portal, export the sessions, look up the official price, multiply in Excel, make a PDF, mail it to the accountant. A small chore, twelve times a year, forever.
For a data scientist that's a bit embarrassing. So I automated it.
The pain: three sources, zero APIs
The numbers live in three places, and none of them were built to talk to each other. The kWh sit in Eve Control, the backoffice of my Alfen charger. The price sits in the CREG "boordtabel", a monthly PDF with the number I need on page 3. In Belgium that CREG price is the reference for paying back home charging of a company car (check the details with your accountant). And the result has to be a document my accounting platform accepts.
The build: one Databricks job, four tasks
On the first of every month at 06:00, a Databricks workflow starts four tasks on a small job cluster. The notebooks come straight from Git, so what runs in production is exactly what's in the repository.
Ingest the sessions. Alfen doesn't offer me an API, so this task starts a headless browser (Playwright), logs in to Eve Control and calls the same CSV export the web app uses. The sessions are merged into a bronze Delta table.
Ingest the price. In parallel, a second task checks the CREG website for new PDFs, stores the originals in a Unity Catalog volume and parses the Flemish household price with pdfplumber.
Transform. Once both are done, every session's kWh is multiplied by the price of its month and aggregated into a silver table. The same task renders the expense note as a PDF with ReportLab.
Deliver. The last task mails the PDF over SMTP, with the password coming from Azure Key Vault. My accounting platform, Okioki, picks it up from the inbox.
The result: €113.98, zero clicks
The first automatic note arrived on 1 October: 46 sessions, 293.75 kWh at 38.80 c€/kWh, so €113.98. A full run takes about 22 minutes, and nobody has to be awake for it.
Why bother with the monthly price at all? Because it moves. The same 293.75 kWh would have cost €201 at the October 2022 peak and €83 at the May 2024 low. A fixed rate agreed once would have been wrong for most of the months since, and looking it up by hand every month is exactly the kind of chore that gets skipped.
The same kWh, a different bill every month
CREG all-in household electricity price for Flanders, in c€/kWh. Hover or use the arrow keys to see what my 293.75 kWh would have cost that month.
View data
| Month | c€/kWh | € for 293.75 kWh |
|---|
Best practices: what I'd do again
Make reruns boring. Every write is a merge on a natural key (charger plus start time). Run the job twice and you get the same answer, and the same amount of money.
Keep the raw stuff. The original CREG PDFs are stored before anything gets parsed. When the layout changes, and it will, I re-parse what I have instead of hoping the old files are still online.
Expect late data. The CREG doesn't publish every month on time. The job then uses the last known price and says so on the document, so nobody has to guess.
Guard the side effects. Computing twice is harmless. Mailing your accountant twice is not. A small log table makes sure every month goes out once.
No API is not a dead end. A headless browser is a fair way in, but it's also the most fragile link in the chain. Treat it that way: alert on failure, and keep credentials in a secret store, never in the notebook.
Why should you care?
You might think, "Bavo, that's a lot of engineering for €113.98 a month." And you're right. But swap the charger for a supplier portal, the CREG for a price index and the expense note for an invoice, and you're looking at a good chunk of a typical back office. Any process where someone copies numbers from one system into another every month is a pipeline waiting to happen.
That's the work I do for clients. Getting data out of systems that were never designed to share it, whether that takes an API, a CSV export or a PDF parser. Setting up the platform underneath, with Databricks, Unity Catalog, deployment from Git and secrets that stay secret. And turning the result into something people actually use: a report, a dashboard, or a document that lands exactly where it's needed.
Got a process that still runs on copy-paste? Let's have a chat, or a coffee. Get in touch.