RSS

Monthly Archives: September 2026

Best Practices for Easy Bulk Data Configuration

A warehouse scene featuring a forklift driver transporting a wooden crate labeled 'DATA'. A woman in a blue jumpsuit stands with a clipboard, checking items off a list. The background shows shelves of boxes and a rules sign on the wall.

Our customer-facing portal and some of our internal tools support bulk configuration via file upload. Over time, our experience has removed a lot of friction from the process making it a great enabler of efficient workflows. A few rules guide our work.

Meet Your Users Where They Are

Of the myriad file formats that can be used for structured data (XML, JSON, CSV, XLSX), some are well-suited for scripts and others are more human-friendly. In business, the universal editor for tabular data is Microsoft Excel. Libraries for C#, Python, and Java (among many other languages) make reading (and usually writing) Excel spreadsheets straightforward.

CSV is an important second option, especially in Linux-heavy environments. The files can be edited in the user’s text editor of choice and tools like CSV Kit allow scripting to manipulate the data. Of course CSV lacks formatting and calculations, but bulk configuration rarely has a need for those.

Rule 1: Support native Excel (.xlsx) first, and offer CSV as a lightweight alternative.

Tell Them What You Expect

It’s a truism that users don’t read documentation. You can write pages describing your upload requirements, but showing is better than telling. Right next to every upload button, provide a one-click download for a sample template with headers your system expects.

Rule 2: Provide a downloadable template at the point of upload.

Ignore What You Don’t Care About

Postel’s Law tells us to be liberal in what we accept. If an uploaded file has an extra column you don’t recognize, is that an error? Not really, you can just ignore it. Don’t assume your system is the only destination for the file. Similarly, ignore case when handling column headers; Username, username, and USERNAME should all map to the same field.

Rule 3: Be permissive with extra data and case variation in headers.

Enforce What You Do Care About

Validating each field in each row is standard practice. But bulk data processing must also have batch-level validation. Before you commit any data from an uploaded file, validate the entire batch, including across rows and against existing data. It’s easy to make sure every value in the batch is unique where necessary. But every value in that column must also be distinct from every value in the database you’re updating. An important edge case arises when swapping two values: the first row uses a value already stored in the database but when the whole batch is processed another row updates the database so the conflict no longer exists. Make sure your upload can handle swapping unique values.

Rule 4: Validate the entire payload against the target state before executing any writes.

Give Specific Feedback

If the user uploads a file of 100 rows, “Update failed” is not useful feedback. “Username ‘asdf$1234’ on row 4 is invalid; usernames cannot include $ or @” tells the user what they did wrong, where, and how to fix it.

Rule 5: Point directly to the row, column, and business logic error.

Give Exhaustive Feedback

It’s incredibly inefficient to stop at the first row with invalid data and make the user iterate until all the problems are fixed. Validation should process all the rows and report all the problems it finds. In the user interface, make it easy to copy that list of problems to another window for reference as the user works through addressing them.

Rule 6: Report every error in the file at once.

Data Rows Start at 2

Programmers may start counting at zero but normal humans do not. Bulk uploads have a header row and data starts at row 2. Provide feedback map internal indices to something natural to the user as seen in their editor.

Rule 7: Translate internal 0-indexed arrays to match the spreadsheet UI.

Frictionless Uploads Build Trust

File uploads are often treated as an unglamorous utility, but for power users, they are a primary interface to your system. Every uninformative error message, rigid schema demand, or broken row index adds unnecessary friction to their workday. By treating bulk upload as a core product feature (one that is forgiving on input and explicit on feedback), you turn a potentially frustrating task into a reliable, efficient workflow.

 
Leave a comment

Posted by on September 22, 2026 in Software techniques

 

Tags: , ,

Small Changes and the Art of Non-Destructive Review

Two landscapers discussing a planting plan in a garden, one holding a blueprint and pointing to a marked spot on the grass, while the other holds a shovel and gives a thumbs up next to a potted plant.

I think we’ve all experienced a system failure that occurs when we are sure nothing changed. But aside from that, it feels true that the smaller the change you make, the better are the chances that it will have no negative repercussions. Having someone review your proposed change improves its chances of success. And small changes are easier to review with confidence than big ones.

When doing file maintenance on a Linux system, I often end up needing to remove some files:

rm <some_pattern>

How sure am I that will do what I intend? Fortunately, listing and removing files in Linux are very close in syntax. I almost invariably precede an rm with a pattern by ls with the same pattern:

ls <some_pattern>

I review the output, convince myself only the intended files are affected, then use the up arrow key to recall the ls, change ls to rm, and press return to execute. I do not retype the pattern!

I can then press up twice and re-execute the ls to verify the files are gone.

Happily, SQL SELECT and DELETE have a similar relationship: the first finds and the second destroys. If I think I want to do:

DELETE FROM my_table WHERE <some_clause>

I generally first do:

SELECT * FROM my_table WHERE <some_clause>

Then, when I’m confident it selects the right rows, I edit the command changing only SELECT * to DELETE and nothing else and then execute it.

It’s less helpful that SQL SELECT and UPDATE have somewhat different forms so if I want to change data, I can’t just write a query and make a simple edit to update the same rows. At least not in an obvious way.

Recently, someone asked me to review their plan to update some data. They offered a fairly complex SELECT that identified the relevant rows and showed the data that needed to be changed. But then they offered a different query to find rows to update. I spent a few minutes looking at the two queries and decided I could not be confident they found and changed the same rows. A common table expression turned out to be the solution.

The query to find the relevant rows ended up something like:

WITH Changes AS (
-- a complex query
)
SELECT t.OldValue, c.NewValue
FROM table AS t
JOIN Changes AS c
ON c.ID = t.ID;

Of course, we could have listed more values from t if they were helpful but this was enough to

  1. Give the developer something safe to iterate on while developing the query logic
  2. Give me a harmless query to review carefully and run as needed to raise my confidence

Once we both felt that the complex logic in the CTE was correct, they had only to change:

SELECT t.OldValue, c.NewValue

to

UPDATE t SET OldValue = cNewValue

That is an almost trivial change that is easy to review. The complex query could have been tens of lines but it didn’t need to be reviewed again because it didn’t change.

They changed it, I reviewed it, they ran it. It did exactly what we wanted.

We often think of database safety in terms of transactions, permissions, or backups. But human-centric safety is just as critical. Good software engineering requires writing code that is easy for a person to read, review, and trust.

If a peer review requires holding two separate queries in your head to verify they target the exact same rows, the process itself invites error. Wrapping the selection logic in a CTE bridges the gap between shell simplicity and SQL structure: it lets you preview the precise impact before swapping out the verb, giving both author and reviewer total confidence.

Small, atomic changes make code easier to write, safer to review, and far less stressful to execute.

 
Leave a comment

Posted by on September 15, 2026 in Software techniques

 

Tags: , ,

Code of Theseus

A historical scene depicting shipbuilders working on a wooden boat at a bustling harbor, with a man sitting contemplatively nearby. The backdrop features a picturesque ancient city with classical architecture and ships in the water.

A recent post on Hack-a-Day about having an AI agent refine a sketch of a script into something production-ready somehow made me think of the Ship of Theseus. If you write the basic algorithm and initial proof of concept but an AI agent optimizes and adds error handling and more, is it still yours?

The script writer had a goal and got something working in short order but, oh, the details! Programs that persist (or crash!), servers that are unavailable or give inconsistent answers. If 90% of programming is error handling, who really wants to write that 90%? (The other 90%, of course, is user interface.)

I’ve written about having an agent do the grunt work and that’s exactly what this developer did. He had the agent fill in error handling and flesh out some features.

This is a good example of how I think these tools work best. The AI handles the implementation, edge cases, tests, and documentation, while the human provides design input and flags anything awkward or that doesn’t fit the intended experience.

This is a workflow I’ve found myself using a lot recently. I sketch something — in a script or simple program, or even as a brief requirements statement — and let the AI build it out. We work together to refine it. Sometimes it’s a realSometimes the agent goes off for a while and comes back with something to review. I look at it, use it a little, offer some critique or suggestion, and we repeat until done.

Come to think of it, this is similar to how I have often worked with interns or junior developers: I give direction and feedback but type a minority of the lines, if any at all. (In process as in product, everything old is new again.)

I’m happy for the help. It lets me concentrate on the big picture. But is the result my work? I’ll leave that to the philosophers.

 
Leave a comment

Posted by on September 8, 2026 in AI, Software techniques

 

Tags: ,