4 ms·
Web Pages via Email – Syntax?
A problem, a solution that is at least
partially complete, some documentation of
the solution, and a question -- is there
more to the problem and more needed for a
complete solution?
The situation and problem: Received some
email that contained a Web page. I.e.,
used Web browser Firefox to go to the Web
site of my email provider, got their Web
page that contained some email sent to me,
in that Web page read the email, and in
that email read a Web page (e.g., from
Goldman-Sachs, but similarly from Macy's,
a university, a symphony orchestra, ...).
So the email (SMTP, simple mail transfer
protocol) made use of MIME (multi-media
internet mail extensions), and the HTML
from Goldman-Sachs was in a MIME Part.
The results looked fine -- Firefox, my
email provider, and Goldman-Sachs all did
fine. But I would like to have the Web
page I received, the one from
Goldman-Sachs, in a file, say, A.HTM, and
that I could give to Firefox and again get
the Goldman-Sachs Web page.
So, just using a text editor, copied the
HTML from the MIME Part to a file A.HTM.
Gave file A.HTM to Firefox and got only a
mess: Some of the text was visible, but
the formatting was a disaster. Congrats
to Firefox for displaying the stuff at
all!
So, Internet Secret 101 A: The HTML data
in the MIME part had some syntax
changes, and to get a file B.HTM that will
display the Web page I was sent (e.g., by
Goldman-Sachs) that will display like it's
supposed to and did, need to undo the
syntax changes!
Why try to get the file B.HTM? Maybe such
a file will prove to be a welcome part of
some future email handling. Maybe.
A guess would be that just using the
capabilities of MIME Parts would permit no
syntax changes, but apparently, instead,
the changes are popular.
So, what are the changes?
Didn't see any documentation so just
looked at some examples, guessed, and saw
two:
(1) The syntax has the lines of HTML bytes
broken (split) at <= 72 characters long
and then an equal character appended to
indicate that should remove the equal and
append the next line to undo the split.
As I recall, this syntax is part of SMTP.
(2) The really popular characters 0-9,
a-z, and A-Z represent themselves, but
many other characters are encoded, e.g.,
a period character can be replaced with
=2D
The 2 and the D are each hexadecimal for
4 bits where the 2 represents bits 0010
and the D represents bits 1101. So, the
2D is 8 bit byte 0010 1101 which is the
period character. There are more details
and characters at
https://www.w3schools.com/charsets/ref_utf_basic_latin.asp
So, to have file B.HTM, have to undo
syntax changes (1) and (2) in file A.HTM.
So, I wrote some code to execute the
undo. On some examples, the one from
Goldman-Sachs and some others, the code
seems to work, got a file B.HTM that
Firefox does display apparently fine, and
changes (1) and (2) seem to be enough.
Okay, Question? I didn't see (1) and (2)
documented so had to guess. So what else
is there to the syntax change beyond (1)
and (2)?
- stop50 2y agoUsually they encode anything except plain text with base64 since it is safe from these problems.
- graycat 2y agoYes, I wondered why they didn't just take the bytes of their HTML, all of them, encode with base64, stuff the result into a MIME Part, send it, and f'get it. Then to read the Web page, just decode the base64, give the result to Firefox, and be done -- no syntax or substitutions. A guess is that base64 would be slower, but the base64 logic is dirt simple. So is the data handling -- just stuff the bytes where want them with no worries about splitting long lines and rejoining them, the fairly tricky logic of finding the =2D strings and replacing them. But the syntax changes seem to be quite popular. Part of all of this, to ease data handling I'm hoping for being able to store each Web page in just one file, a one file HTML, especially for whatever I receive via email.