Why is the .NET Framework StreamReader / Writer default for UTF8 encoding?
UTF-8 will work with any ASCII document and is generally more compact than UTF-16, but it still covers all of Unicode. I would say that UTF-8 is much more common than UTF-16. This is also the default for XML (when there is no BOM and no explicit encoding specified).
Why do you think it would be better to use UTF-16 by default? (This is what it Encoding.Unicode
is.)
EDIT: I suspect you are confused about what UTF-8 can handle. This page describes it pretty clearly, including how any particular Unicode character is encoded. This is a variable width encoding, but it covers all Unicode.
a source to share
Like everyone else, UTF-8 is the Unicode encoding standard. UTF-8 uses a variable number of bytes to encode all Unicode characters.
All ASCII characters are presented as-is, so ASCII files can be read further ado. Once a byte in the stream has 8th bit (most significant bit,> 127), it forces the reader to combine it with the next byte until it is <128. The combination then counts as 1 character.
LATIN-1 (ANSII) has characters that are encoded using two characters: for example, é is encoded as e and '. The length ('é') is therefore 2.
Windows uses UTF-16 internally, which limits the encoded characters to 64K, which is by no means all Unicde characters. UTF-32 currently allows all characters, but is also artificially limited. And both are not ASCII compatible since they have leading zeros:
A = ASCII h41 = UTF-8 h41 = UTF-16 h0041 = UTF-32 h00000041
There are also small and large endian encodings:
A = UTF-16 big endian h0041 = UTF-16 little endian h4100
Imagine you are using UTF16 or UTF32 to save your files. They would be (for text files) two or four groups in size compared to ASCII and UTF-8 (UTF-8 if only ascii characters are used). UTF-8 not only allows all the characters in the Unicode standard to be used, even for future enhancements, but also effectively saves space.
Typically, the first two bytes of the file, specifier, or order byte marker tell you which encoding standard is being used. If you skip it, XML and StreamRedaer both use UTF-8, as you learned. This again results in ASCII files lacking a BOM and therefore reading correctly in most cases. This may not be true for files using all LATIN-1.
a source to share