Skip to content
A scene from Ireland

Beyond Spreadsheets

It is time that so-called "Office suites" got beyond the spreadsheet.

Spread sheets are a wonderful and powerful idea. I have put them to good use several times. But they are tied to an idea that is no longer relevant - the idea from which they got their name.

A spread sheet is basically an infinitely (or atleast indefinately) large two dimensional table of cells. Each cell can do calculations based on values in other cells. This is great for doing calculations on columns or rows or whole tables. This works really well. There is just one problem.

I've never wanted to have an indefinately large table.

When I'm working with some data in a spreadsheet I might have a few numbers, or some rows or a whole table or two. But each of these have a well defined size, they aren't indefinate. It is true that I might want to add rows, and occasionally even columns, to a table, but when I do, I want to actually add them. I don't really want them to already be there, as they might well be in the wrong place.

But a spread sheet doesn't give you a few cell, rows, and the odd table. It gives you a virtually infinite sheet that you can put stuff in. And that is what people (well, myself at least) tend to do. There is room for a couple of tables side by side, a few random values and calculations and whatever. But the point is that it is all random. There is no right place to put things, so things go anywhere. An then if you want to add a row into the middle of a table, you end up inserting it in the middle of all the side-by-side tables that you might have. It was when this first happened to me that I realised something was wrong.

Many spread sheets today have multiple sheets open at once, with easy linking between them. This removes my immediate problem if not being able to insert a row in just one table - you have a separate sheet for each table. But still I feel something is wrong.

As I understand it, the name "spread sheet" can from a pencil-and-paper tool of the same name that modern spread sheets are based on. A large sheet of paper with a pre-ruled grid would be spread out, and various tables of data would be written onto it, presumably well spaced out. Then various calculations on this data would be performed and the sheet provided place to put the answers where they could easily be kept track of. As it is not possible to simply add columns on a sheet of paper, having lots of columns and rows pre-ruled help out a lot.

But that simply doesn't apply with a computer based spread sheet. It is trivial can very common to add rows and columns to a pre-existing table. So there is absolutely no need to have thousands of rows and columns pre-existing (even if they are only virtually existing).

In my mind, the way forward for spread sheets is to discard them as separate entities and simply add a "cell" concept to a standard word processor.

Word processors already have tables which can be created and extended at will, and which have well defined column headers and such. If one can create a table and embed a spread-sheet-like cell in every cell of the table, then you have the calculating power of a spread sheet right there with the formatting power of a word processor.

Ofcourse calculation cells don't have to appear in tables. They can appear anywhere in the document that a calculation is performed. Instead of creating a link to a separate spreadsheet when the text says that the mean of some set of values is "X", the calculation can be right there in the document.

Once we are liberated from the "infinite table" model of calculation there are other changes that become obvious. One this about spreadsheets that I don't like is that when you have a summary column - say column X is the sum of columns B through G, then you have to place that calculation in every single cell, rather than just once for the whole column. There are almost always very simply tools for duplication the cell suitably and it isn't hard. But it feels untidy that you have to.

The reason that we have to is obvious - normally you don't want the whole infinite column to have the same calculation. You probably want a separate heading in the first cell of the column, and you possibly want some other sort of column-summary just beyond the last cell of the column.

With a word processor (atleast a real word processor like TeX - I'm not sure exactly how WYSIWYG word processors work, their tables are almost as ad hoc as a spread sheet) There is often a default style for all cells which can be over-ridden for headers and footers. So the default style for a column could include a cell calculation which automatically then appears in every row that is created where the default is not explicitly over-ridden.

I would like to then take this model one step further and incorporate some simple database concepts into it. A table that has data explicitly entered into it should be treated much like a simple database. Columns in a table that are raw data are stored in the database. Columns in a table that are calculations are evaulated on the fly.

Then you could create a separate table in the word processor which has columns selected fromn the database, and instead of explicitly listing rows, it could be given a row-rule - a way to extract rows from the database such as "Column X is less than 42". This would make it possible to create reports based on subsets of "spread sheet" data without selectively hiding column and rows - another pet peeve of mine.

Doing this would smooth the current rather artificial distinction between a spread sheet and a database, and the distintion between a word processor and a report generator.

Ultimately I would like to have just one tool that could do arbitrary formatting, arbitrary calculation, and can access tabular data either in a simple file or behind some sort of SQL interface.

Then you just need to add an obnoxios background and use a large font, and the same tool becomes your presentation manager.