Skip to content

Comment on A collection of high performance c-string transformations

Comments

Does it have UTF / other format support? I really have no use if it doesn't, and I don't see anything mentioning it on the first page.

This was my question, so I looked at the source. It refers to a page on algorithms that claims that it is 8-bit clean, i.e., it preserves non-ascii characters but does not modify them. I then compiled a test program to verify this, because the bit twiddling in the source code looks a little nuanced. It IS 8-bit clean.

Although, if you're looking for something that converts non-ascii characters to their lower case equivalents, just use LibICU. This library looks nice for decoding things like HTTP which have plenty of ascii-only case insensitivity.

Since it is doing bit manipulation, then probably not. I mean, the toupper function is treating (4) 8 bit chars as one 32 bit uint. Then doing math... so, I assume that it will miss the finer points of I18N.

However, the base64/85 en/decoders should work just fine.

And in the same vein as the nginx optimizations suggested in a comment yesterday, it fails on arcitectures that don't support unaligned access.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.