Skip to content

Comment on Xerox scanners and photocopiers randomly alter numbers in scanned documents

Comments

I can't quite see the reason why you would lossily compress something when your machine's purpose is to duplicate things.

Anyone got a reasonable reason for doing this?

In the good old days of analog copiers this would be impossible - the scanner send the light through a system of mirrors to the drum, the drum gets static charged, the toner is pulled on the charged parts and gets transferred to the transfer belt, here the paper has the opposite charge and pulls the toner off of the transfer belt, goes through the fusing unit and here is the toner 'burned' to the paper. End of Story

On a modern copier the scanner transfers the data first to RAM and than usually to a hard disk (the most of the people do not even know that the "copy machine" has one and saves the scanned stuff to it). From that hard disk the data where transmitted via laser to the drum

Tadaaa - you have the reason for having data be compressed on a modern copier.

Yup, and those old analog copiers - good ones at least - had beautiful crisp output. The resolution was good enough to reproduce printing dots so they could even duplicate photos from books. Continuous tone of an analog photograph didn't work as well. They sure were expensive though.

I might have missed something, but my reading is that the article doesn't state or imply this happens with regular photocopies, only with scans to PDF.

Others have pointed out a credible explanation: to have the document take less space on their hard disk.

However, it does not have to be compression, per se. Modern copiers want to correct all kinds of errors such as creases and staples. They also want to optimize the colors. To do that, they have logic for detecting what areas of the page are full-color and which are black and white, which are half-tone printed, which are text, line art, photograph, whether the paper might have aged, etc.

I don't know what tricks they use, but I do not rule out that they will replace 'looks somewhat dirty' patches with an 'obviously higher quality version' of them, and use too aggressive parameters in some of those heuristics.

Well we have 14TiB of financial documents archived on our kit. There is no way we even would consider such compression!!!

The whole thing is dangerous and wholly illogical.

This is akin to a crappy crime flick where someone hits the "enhance!" button on a CCTV still a few times and gets to see the dirt on the guy's teeth.

In this case, the computer decides the guy is female and has no teeth.

IIRC when security cameras moved from individual frame compression algorithms like M-JPEG to modern codecs which could sometimes replace small movement in background with still image if there is a bigger change in foreground there were news reports about some problems with investigations.

If you're scanning a long document to a PDF, compression makes a lot of sense. It's the difference between being able to email the PDF as an attachment and having to find a place to put the file online.

Exactly, and that is why there should be a compression step on the code path that handles the paper -> pdf case. This doesn't make any sense in a paper -> paper case, however, as any electronic version of the image will only be stored internally, for a very brief time.

It is amazing the things people expect to be emailed. Can you email me the MRI scan? It's over 1500 images. Can you email it?

It's amazing to me the engineers who refuse to update their worldviews about normal people's mental models for "sending data" that still get amazed by this.

The size of the files or the number of them are totally irrelevant.

Normal people seem to get that it's considerably harder to ship a barn than a letter, and that if you want to move a barn you use a specialty service rather than the post office.

The size and number of files are and should be totally relevant even to "normal" people. When someone asks for something in e-mail, it's perfectly reasonable to say "no, it's much too big" and expect them to understand.

But we're not dealing with barns or letters or any physical object, we're dealing with abstract systems where the physics are much more flexible and changeable. It's important to change our computer systems to work for us, rather than attempting to change people to adapt to the computer systems. We should discard those systems that can not adapt to humanity, as they are of little worth in the long run.

When somebody says, "can you email it to me" they mean, grant access to the data via their centralized messaging system, their email. There are many ways to make that happen, one of which is an attachment, another of which is linking to the content, but the key is to make sure that it's low friction and takes very little time or clicks to get access from the email.

There are some pieces of content for which it's entirely impractical to "grant access via e-mail". On occasion people ask to be e-mailed extremely large blocks of data, where it would literally be faster to burn it to a pile of DVDs and then FedEx them than to upload-and-then-download the data. Depending on the size of the medical images mentioned in a previous post, that might actually be the case in that circumstance.

It's a failure of technology when it's difficult to send ordinary-sized files like a few photos or a couple pages of documents. But it's a failure of people when they don't recognize the possibility that some types of data (video, large numbers of images, scientific research data, whole databases) simply can't be sent quickly, yet they fail to plan ahead to gain access. (I've also entirely skirted the issue of "some data should have its access restricted physically"...)

But it's a failure of people when they don't recognize the possibility that some types of data (video, large numbers of images, scientific research data, whole databases) simply can't be sent quickly, yet they fail to plan ahead to gain access. (I've also entirely skirted the issue of "some data should have its access restricted physically"...)

I disagree vehemently with that attitude, and I have to deal with it everyday. In my field, >50% of the data we receive is transferred by overnight courier of hard drives due to quantity of data. It's a crappy attitude to blame people for having to learn that, and in an ideal world we'd share it via access granted by email. People should not be blamed for not understanding that, our infrastructure should be blamed for not supporting 10Gb everywhere, and cheap access to 40Gb+ on long-distance connections.

Nothing is helped by blaming people, and relationships can be harmed by doing that. But we can change the technology.

As a side note, DVDs? Really? They're incredibly slow at data transfer once you have them in hand, the tiny size of a DVD requires tricky archive spanning methods, and optical discs are flaky technology all around. Hard drives or LTO-5/6 all the way.

I think most people would understand if you asked them how long it would take to download a million large photos, given that one large photo often takes several seconds to complete. They'd realize that this might be a slow process.

Incidentally, it's not just about transfer speed. Sometimes people ask if you can e-mail something that has never been put on a computer, and would take weeks or months to scan in. Or sometimes they ask for access to information when access is very slow to set up due to security or privacy considerations. Or sometimes they ask for access to something that the boss needs to physically sign off on, after the boss has gone home for the day. This is only a problem if they've decided it's urgent to have it, and simply haven't thought ahead about how it might not necessarily be possible to get instant access to every piece of information that ever existed.

We can change technology. But we also need to retain the mindset of arranging access beforehand. It's not about "blaming", it's simply about encouraging people to understand what they're asking for and to make sure they get the access they need before they need it.

[As an aside, DVDs are just an example of "sometimes it's really freaking slow to download data" that somebody like my mom would get. An alternative way to phrase it would be "downloading that would be so slow, it'd be better to just have your friend bring her laptop over." I certainly don't intend to suggest a new industry standard.]

Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway

Andrew Tanenbaum

It would be helpful if file managers gave better cues as to file size. A barn is obviously different to a letter, but a one byte file is normally given the same icon as a one terabyte file.

I think we can all agree that e-mails should have finite size - it's not a very good protocol for transferring multi-gigabyte files, for sure! Where we would disagree is where that limit should be drawn.

I've seen systems in this day and age that fail in the face of e-mails as small as 5 megabytes (e.g. Yahoo Popgate) which IMHO is far too low - but evidently some sysadmins disagree with me!

Email size is a technical issue that shouldn't be limiting (or even visible) to the end user. If an end user wants to send a multigigabyte file to another user's email address - why not? The email client could launch a background upload process and email a link to get that file by, say, bittorrent... Some protocol extensions and software support would be needed, but that can be done and, as users need it, probabpy should be done.

Oh, the protocols are already in place; consider RFC2017.

Sending a link to something (even wrapped in a nice ui and container) has pretty different semantics from actually sending the something, though.

Opinions run the gamut. I'm firmly in the "email should be plain text" camp but realized that battle was lost long ago.

Just to make it clear, I'm not an engineer (I'd like to be good enough to be considered one, but that's a long way off). I'm more of an enthusiastic amateur who knows enough to badly break things. And the people who ask most frequently ask are ortho surgeons with a patient asleep on the table. Anyone who waits until that late in the piece then asked for a 1.5gig email becuase they weren't organised enough to sort out access to images in the 3 month lead up to an operation is not a normal person.

We developed an easy way to email and collaboratively view CT or MRI studies. http://www.claripacs.com.

Thanks, I'll be looking into that.

Emailing the file or emailing a link to the file is just as useful. As long as you have an encrypted document sharing capability you should be able to say sure to just about anything.

PS: Granted that's assuming fast networks. For 80+ Gig VM's sending a removable drive is often faster.

For a lot of end users, if it's not a proper attachment no end of grief is caused.

And encryption and file hosting causes more hassle in enterprise environments with unforgiving compliance policies.

The article has been updated with the probable cause for this error.

Cheaper components, maybe? (If it lets them get by with less memory for example.)

Good point. Looking at a product page (http://www.office.xerox.com/multifunction-printer/color-mult...), I see that the first model mentioned is multifunction, it can "Copy, email, fax, print, [and] scan".

So it sounds like there's one code path and it's seriously broken. I looked at the first settings page, and while it's in German I can see it's 200 DPI. There's no excuse for default lossy compression when you're at 200 DPI and doing office sized paper. We didn't do that in 1991, we got CCITT Group 4 lossless compression of around 50KB per image plus or more generally minus for 8.5x11 inch paper, although we did do thinks like noise reduction and straightening documents (that makes them compress better, among other things).

CCITT Group 4 also known as Modified Modified Read is in no way something you'd be wanting to use now ever. I'm telling you why MMR sucks hard:

1. It's monochrome. No greyscale, no color. This works for text and lines, but nothing else. No big surprise, it was designed for Fax. But this makes CCITT G3 and G4 lossy.

2. It has no defined endianess. This adds another fault risk which you won't see coming as long as you're working on an isolated platform but can hit you in the nuts when you change hardware or software.

3. The data does not contain resolution or dimensional information, as well as no information about endianess This means that you have to rely on a container providing these informations. It could be TIFF, it could be PDF, it could be something an intern coded during coffee break. This is good on one hand, but evil on the other. Software is sold, saying CCITT G4 compression (a standard, after all) is used, while the data can be embedded in proprietary containers.

4. It's a 2D compression, meaning the compression is applied on a matrix of binary pixel data. As the standard does not specify the dimensions, you depend on another image container like TIFF to provide information. Because G4 removed EOL markers, there is no way to reconstruct image dimensions from the compressed data alone.

5. It's not exactly fault tolerant. Transmission errors can influence larger areas of the image up to making the picture totally unreadable. Flipped bits are not too critical, missing bits are, due to the 2D compression.

There are many excellent, fault tolerant, standardized Image formats ready to use for document processing and archiving, CCITT G4 isn't exactly one of them.

Errr, I didn't communicate clearly.

What I meant to say was that CCITT Group IV gave acceptable sizes for early-'90s computing power, CPU and disk, and something at or better than its level of lossless compression today should be even more acceptable.

And in light of this screwup, I suspect we'd agree that Xerox would have been better off to use lossless (well, after the scanning, as you point out, but then again no one was willing to pay for color) CCITT Group IV than overly clever lossy JBIG2.

"It could be TIFF, it could be PDF, it could be something an intern coded during coffee break."

It could be something a journeyman software engineer edging to expect coded in a Saturday afternoon in a very fast paced project; for me, 3 weeks on the "engine". And, oh my, I can't remember encoding endiness, except of course for the leading TIFF bytes. But I had a guy who knew this cold telling me what to do, he was the one who debugged all our raw compressed data problems bit by bit. And, yeah, it was an "Intel" little endian TIFF, and I think I recall the Kodak Powescans produced that (600 pound monsters that could scan 18 inches per second at 200 DPI).

Hmmm, at least back then, "TIFF" was the selling point, and, oh yeah, it's Group IV compressed (except of course when it wasn't, we once dealt with some weird enhanced Group III).

Of course they would have been better off with T.6, as Group 4 at least did not modify the image content. However especially with TIFF there are/were countless implementations of viewers, components, libraries and every single one of them had their own habits. Some would not regard endianess, some would assume payload endianess is the same as the TIFF, some did respect the Tag for byte order specific to the image. When I coded my first TIFF Library, I was around 14 and the most troublesome part of doing it was keeping myself from bashing my head against the next available wall due to stupidity of other people who thought interpreting a standard according to their wishes was ok, because there'd never be someone trying to display the images with a viewer different from theirs.

I don't know how deep you have dived into TIFF, but maybe you remember the TIFF6 Standard way of embedding JPEG. It was the biggest pain in the ass imaginable, having to parse JPEG files, splitting them and packaging it into different TIFF Tags. Before TTN2 and easy embedding of JPEG Images, everyone invented their own way of avoiding the standard. Some defined their own compression type, some used the standard compression type, but used it in a nonstandard way, ah, I'm starting to lose my hair again ;-)

Not that deeply, I only did B&W document imaging, and I think the last time I worked on TIFF headers and tags was in 1992, so it was almost certainly the 5.0 standard, 6.0 came out in that year.

And yeah, it was a mess; we mostly did the best we could and made sure the ones we generated worked for our customer's reader(s). Although I don't remember any big problems with people reading the ones we produced.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.