HTML Essentials
SAP Implementation Consulting & ERP - Startupistan Germany · Block I · study notes for revision.
Where HTML sits: the three languages of the web
Section titled “Where HTML sits: the three languages of the web”Almost every page on the web is built from three languages, and each has exactly one job. Getting this split clear first makes everything below easier - HTML never tries to look good, and that is on purpose.
| Language | Its one job | Says things like |
|---|---|---|
| HTML | Structure & content - what each thing is | ”this is a heading”, “this is a list” |
| CSS | Appearance - what it looks like | ”this heading is dark green, in this font” |
| JavaScript | Behavior - what it does | ”when this button is clicked, do X” |
This chapter is only the first row. The web deliberately keeps them separate - most other software mashes structure, looks, and logic into one codebase, but the web splits them so a design can change without touching content, and so machines that only read structure (search engines, screen readers) can understand a page without running any code.
Two words worth defining now, because they run through everything:
- Browser - the program that reads these three languages and paints the page (Chrome, Firefox, Edge, Safari). It is the universal client: written once, a page works on every device with nothing to install.
- Semantic - “carrying meaning”. A semantic element says what its content is; a non-semantic one says nothing. This idea is the thread through the whole chapter.
What HTML actually is
Section titled “What HTML actually is”HTML = HyperText Markup Language. Break the name apart:
- HyperText - text that contains links to other text. Links are what turned separate documents into a connected web.
- Markup - annotating content with labels that say what each piece is. The labels are not the content; they are information about the content.
- Language - it has fixed rules and a fixed vocabulary that every browser agrees on, which is why the same page renders the same everywhere.
The single most important idea: HTML says what content is, not what it looks like. When I write <h1>My Bakery</h1>, I am not saying “make this big” - I am saying “this is the most important heading on the page”. The browser chooses to show it big by default, and CSS can later make it any size while it stays the most important heading.
The same words with three different labels mean three different things to a machine:
<h1>Fresh bread every morning</h1> <!-- the page's main heading --><p>Fresh bread every morning</p> <!-- a paragraph of text --><li>Fresh bread every morning</li> <!-- one item in a list -->Identical text. A search engine weighs the h1 words far more than the p words, and a screen reader (software that reads a page aloud for people who cannot see it) treats each completely differently. Honest labelling is not pedantry - it is how the page stays usable for every reader, human and machine.
Anatomy of an HTML document
Section titled “Anatomy of an HTML document”Every proper HTML page - from a five-line placeholder to a global newspaper - has the same four-part skeleton, in this order. Here it is, annotated:
<!DOCTYPE html> <!-- 1. Declaration: "read this as modern HTML" --><html lang="en"> <!-- 2. Root element, wraps everything; lang = content language --> <head> <!-- 3. Info ABOUT the page - not shown on the page --> <meta charset="UTF-8"> <!-- which alphabet the file uses (UTF-8 = all languages) --> <meta name="viewport" content="width=device-width, initial-scale=1.0"> <title>My Bakery</title> <!-- shows in the browser TAB, bookmarks, search results --> </head> <body> <!-- 4. Everything the visitor actually SEES --> <h1>My Bakery</h1> <p>Fresh bread, honest coffee.</p> </body></html>The four parts, top to bottom:
| Part | What it is | Key point |
|---|---|---|
<!DOCTYPE html> | A declaration, not an element | Always line 1, always this exact spelling, never has a closing partner |
<html> | The root that wraps everything | lang="en" names the language - screen readers & translation rely on it |
<head> | Info about the page (metadata) | Nothing here is drawn on the page itself |
<body> | Everything the visitor sees | Headings, paragraphs, images, links all live here |
Elements, tags, and attributes
Section titled “Elements, tags, and attributes”Three words appear on every page of this topic. Developers say them loosely; error messages and documentation use them precisely, so I learn them precisely.
- Element - a complete piece of the page: the markup plus its content.
<p>Fresh bread</p>is one element. - Tag - the bracketed marker that opens or closes an element.
<p>opens,</p>closes; the slash is what makes it a closer. - Attribute - a
name="value"pair inside an opening tag that configures the element. In<html lang="en">, the name islangand the value is"en".
So a typical element reads: opening tag (maybe with attributes) → content → closing tag.
<p lang="de">Frisches Brot, ehrlicher Kaffee.</p>One element, two tags, one attribute (lang="de", marking just this paragraph as German), content in the middle.
Void elements carry no content, so there is nothing to close - no closing tag. The two I meet constantly are <meta> and <img>; everything they need is said in attributes. Void elements are the exception; when in doubt, an element closes.
Nesting - elements contain elements (the body contains the h1; the html contains everything). One rule governs it: the element opened last closes first. Boxes inside boxes, never overlapping rings.
<!-- WRONG: strong opened last but p closes first - they overlap --><p>Our bread is <strong>baked fresh</p></strong>
<!-- RIGHT: close in reverse order of opening --><p>Our bread is <strong>baked fresh</strong></p>Indentation (two spaces per level) is purely for humans - the browser ignores it completely. I indent so future-me can read the nesting.
The tag reference table
Section titled “The tag reference table”The most-used tags, grouped. This is the working vocabulary for the whole chapter - I don’t memorise it, I recognise it.
Headings & text
| Tag | Purpose | Example |
|---|---|---|
<h1>-<h6> | Section titles, ranked by importance (1 = top) | <h1>My Bakery</h1> |
<p> | One paragraph of text | <p>Baked every morning.</p> |
<strong> | Content that is seriously important | Closes <strong>Thursday</strong>. |
<em> | Vocal stress that can change meaning | We bake <em>everything</em>. |
<br> | A single line break (void) | Line one<br>Line two |
Lists
| Tag | Purpose | Example |
|---|---|---|
<ul> | Unordered list - sequence doesn’t matter | <ul><li>Milk</li></ul> |
<ol> | Ordered list - sequence is the point | <ol><li>Preheat</li></ol> |
<li> | One item in either kind of list | <li>Croissant</li> |
Links & images
| Tag | Purpose | Example |
|---|---|---|
<a> | A link, pointing wherever href says | <a href="menu.html">Menu</a> |
<img> | An image (void), from its src | <img src="photo.png" alt="..."> |
Structure / semantic containers
| Tag | Purpose | Example |
|---|---|---|
<header> | Intro area - site name & navigation | <header>…</header> |
<nav> | The navigation links themselves | <nav>…</nav> |
<main> | The unique content of this page (one per page) | <main>…</main> |
<section> | One thematic group, usually with its own heading | <section>…</section> |
<article> | A self-contained piece (a post, a card) | <article>…</article> |
<footer> | Closing area - small print, secondary info | <footer>…</footer> |
<div> | Block container with no meaning (for grouping) | <div>…</div> |
<span> | Inline container with no meaning (few words) | Price: <span>5.80</span> |
Tables & forms (full detail in their own sections below)
| Tag | Purpose | Example |
|---|---|---|
<table> <tr> <th> <td> | Grid / row / header cell / data cell | see Tables section |
<form> | Container for one set of questions + submit | <form>…</form> |
<label> <input> | Field name / the field itself | <label for="x">… |
Headings and text semantics
Section titled “Headings and text semantics”Headings are an outline, not a font-size picker
Section titled “Headings are an outline, not a font-size picker”h1-h6 are ranks of importance, not sizes. Together they form the page’s outline, like a book’s chapters and sections. Two rules keep the outline honest:
- One
<h1>per page - it names what the whole page is about. Twoh1s = two claims to the same throne. - Don’t skip levels - after an
h2, go toh3when you go deeper, not straight toh4.
Browsers render higher levels bigger by default, which tempts beginners to pick a heading by its size. Resist it - CSS will make any heading any size later. The level is a statement about meaning that outlives every design change.
Paragraphs and collapsing whitespace
Section titled “Paragraphs and collapsing whitespace”<p> marks one paragraph. The surprise everyone hits once: the browser collapses whitespace. Any run of spaces, tabs, and line breaks in the source becomes a single space on the page. Pressing Enter five times changes nothing - only a new <p> creates a new paragraph. This is freeing: I can format the source for human readability and the rendered page is unaffected.
strong vs em - meaning, not bold and italic
Section titled “strong vs em - meaning, not bold and italic”Both look like styling switches (browsers show strong bold and em italic by default), but they carry meaning:
<strong>- this content is seriously important (a warning, a deadline, a key fact).<em>- this word is stressed, the way my voice would stress it aloud. “We bake<em>everything</em>ourselves” vs “We bake everything<em>ourselves</em>” stress different claims with the same words.
Three elements, and the choice between them is a meaning choice like everything in HTML:
<ul>- unordered list (sequence doesn’t matter): a shopping list.<ol>- ordered list (sequence is the point): recipe steps.<li>- one item, used inside both.
<ul> <li>Butter croissant</li> <li>Cinnamon roll</li></ul>Structural rule: the only direct children a ul or ol may have are li elements. Text and other elements go inside the list items, never loose between them.
Nested lists build categories - a list item can hold a whole list, which is how a menu with categories is shaped:
<ul> <li>Breads <ul> <li>Country sourdough</li> <li>Baguette</li> </ul> </li> <li>Pastries <ul> <li>Butter croissant</li> </ul> </li></ul>The inner <ul> lives inside its category’s <li>, before that li closes - the nesting rule doing real work. Don’t hand-type “1.” and “2.” inside paragraphs: a screen reader announces a real list (“list, four items”) and lets users skip it; typed numbers give none of that.
The H in HTML is HyperText - links are what make a web a web.
Reading a URL
Section titled “Reading a URL”A URL (Uniform Resource Locator) is the address of a resource. Every one decomposes into three parts:
https- the scheme: which protocol to use (here, the secure web protocol).example.github.io- the domain: which server to ask./my-site/- the path: which resource on that server.
The <a> element
Section titled “The <a> element”<a href="menu.html">See our menu</a>The text between the tags is what the visitor clicks - write text that says where it leads (“See our menu”, not “click here”). Screen-reader users often jump link to link hearing only the link texts, and a page of “click here, click here” tells them nothing.
Two kinds of address in href:
- Relative - points from the current file to another file in my own project:
menu.html,images/hero.png. No scheme, no domain. Relative links keep working everywhere the project goes (local preview, live server, a teammate’s copy). - Absolute - a full URL with scheme and domain:
https://www.example.com. Use it only for destinations outside my project. Hard-coding my own live URL into internal links is the classic self-inflicted wound: the site then breaks everywhere except that one address.
Images
Section titled “Images”<img src="images/hero.png" alt="The bakery counter with fresh loaves on wooden shelves."><img> is a void element - no closing tag, because its “content” is the picture file its src points at. Everything is said in attributes.
src- the path to the file. Convention: keep images in animages/folder in the project and reach them with a relative path. The image file is committed, reviewed, and released with the markup that uses it - always in the same commit, because one without the other is a broken state.alt- the text a visitor receives instead of the image: read aloud by screen readers, shown when the file fails to load, read by search engines. Every image gets an alt sentence - a full sentence a person could hear and still follow the page, not a keyword and not the file name.
Page structure and containers
Section titled “Page structure and containers”Five semantic elements name the areas almost every page has. Each is a meaning label - like strong, but for a whole area:
<body> <header> <nav> ... </nav> </header> <main> <section> ... </section> </main> <footer> ... </footer></body><header>- intro area, usually site name and navigation.<nav>- the navigation itself, a group of links.<main>- the unique content of this page (one per page).<section>- one thematic group within the content, usually with a heading.<footer>- closing area, small print and secondary info.
These barely change how anything looks. What they change is what the page means: screen readers announce them as landmarks and let users jump straight to main or nav, skipping repetition. Search engines and reader modes lean on them too, and accessible structure is a legal matter for many public sites - it’s the cheapest compliance work there is, free at writing time.
The two meaning-free containers, for when I need to group things purely to style them later:
<div>- a block container that says nothing about its content.<span>- an inline container that says nothing, for a few words inside text.
<p>Whole loaf: <span>5.80</span></p>Block vs inline
Section titled “Block vs inline”Every element renders one of two default ways. I’ve been seeing this all along - here’s its name:
| Block-level | Inline | |
|---|---|---|
| Line | Starts on a new line | Flows within a line |
| Width | Takes the full available width | Takes only the space its content needs |
| Examples | headings, p, lists, div, all 5 structure elements | strong, em, a, span |
| Behaviour | Stack vertically, each its own row | Live inside lines, never break them |
Block elements pile up like stacked bars; inline elements sit in the flow of text. This is the default behaviour of every element, visible before a single line of CSS - later, CSS’s display property can change it, which is the doorway to layout.
A site with more than one page
Section titled “A site with more than one page”A small site is just several complete HTML documents sitting side by side in the project root, next to the images/ folder:
index.html- the home page. It keeps this special name because a web server, when asked for a folder rather than a file, answers with that folder’sindex.html. That’s why a live address likeexample.github.io/my-site/needs no file name.menu.html,order.html- other pages, named descriptively. Only the default page gets the special name.
Each file is a complete document: full anatomy, its own <title> (so three open tabs are tellable apart), its own single <h1>.
One nav, on every page - for a small site, that means the same markup copied into each file:
<nav> <a href="index.html">Home</a> <a href="menu.html">Menu</a> <a href="order.html">Order</a></nav>Yes, copies - when the nav changes I update all of them. Feel that small pain: repeating identical markup is a real problem with real solutions later (JavaScript), but for a few pages honest repetition is the right call. Every page links to all pages, itself included, so the nav stays stable.
Tables for data
Section titled “Tables for data”A table holds rows and columns of related data - and nothing else. (Before CSS could do layout, developers abused tables to position whole pages; that practice is dead. Tables hold data; layout belongs to CSS.)
| Tag | Purpose |
|---|---|
<table> | The grid of data |
<thead> / <tbody> | Wraps the header row / the data rows |
<tr> | One table row |
<th> | A header cell (names its column or row) |
<td> | One data cell |
<table> <thead> <tr> <th>Day</th> <th>Hours</th> </tr> </thead> <tbody> <tr> <td>Saturday</td> <td>7:00 to 14:00</td> </tr> <tr> <td>Sunday</td> <td>Closed</td> </tr> </tbody></table>thead and tbody aren’t decoration: screen readers use them to announce each header with the data it belongs to, and CSS uses them to target header vs body rows separately (for example, striping every second data row). Data tables are the correct, accessible element for data - nothing else does their job.
Web forms
Section titled “Web forms”A web form is a section of a page that collects input from the visitor and sends it somewhere for processing. Two halves: the collecting (HTML’s job, and the whole of this section) and the sending (explained conceptually here, made real in later courses).
Forms are everywhere once you notice: the search box on every site, logins, checkouts, newsletter signups, a pull request’s title and description. Forms are how the web listens - the one mechanism by which pages stop being read-only.
Building a form
Section titled “Building a form”The pieces, from the outside in:
| Tag | Purpose |
|---|---|
<form> | Container for one set of questions + its submit control |
<fieldset> | One group of related questions |
<legend> | The visible title of a fieldset |
<label> | The visible name of one input, wired to it by for |
<input> | One question; its type decides what kind (void element) |
<textarea> | Multi-line free text |
<select> + <option> | Choose from a fixed list |
<button type="submit"> | Sends the form |
Four attributes that matter on inputs:
type- what kind of question. Common:text(free text),email(browser checks for email shape),radio(choose exactly one of a group),checkbox(choose any number). Radio for one, checkbox for many. Radios sharing onenameform one choose-one group; giving one thecheckedattribute preselects a default.name- the key the value travels under (customer-name=Sarah). An input with nonamesubmits nothing, silently - the field looks perfect and its data never leaves. This is the most important attribute to never forget.required- the browser refuses to submit while the field is empty, with its own built-in message, no code needed.placeholder- faint example text inside an empty field. An example, never a label - it vanishes the moment typing starts.
Here is a full, copy-pasteable form showing every rule at once:
<form action="#" method="get">
<fieldset> <legend>Who is picking up</legend>
<label for="customer-name">Your name</label> <input type="text" id="customer-name" name="customer-name" required placeholder="e.g. Alex Example">
<label for="customer-email">Email</label> <input type="email" id="customer-email" name="customer-email" required> </fieldset>
<fieldset> <legend>Pickup day</legend>
<input type="radio" id="day-sat" name="pickup-day" value="saturday" checked> <label for="day-sat">Saturday</label>
<input type="radio" id="day-sun" name="pickup-day" value="sunday"> <label for="day-sun">Sunday</label>
<label for="pickup-time">Pickup time</label> <select id="pickup-time" name="pickup-time"> <option value="morning">Morning</option> <option value="afternoon">Afternoon</option> </select> </fieldset>
<fieldset> <legend>Your order</legend>
<input type="checkbox" id="item-loaf" name="items" value="loaf"> <label for="item-loaf">Sourdough loaf</label>
<input type="checkbox" id="item-croissant" name="items" value="croissant"> <label for="item-croissant">Croissant</label>
<label for="notes">Notes for the bakers</label> <textarea id="notes" name="notes"></textarea> </fieldset>
<button type="submit">Place pre-order</button>
</form>Read the radio group carefully: both share name="pickup-day" so they exclude each other, each has a distinct id, each has its own wired label, and Saturday is preselected with checked.
What happens on submit
Section titled “What happens on submit”Two attributes on <form> describe the sending:
action- the address the data is sent to.method- how it travels:getorpost.
With action="#" (this same page) and method="get" and no server at all, I can watch a submission happen: fill the form, press submit, and read the address bar. The URL now carries every named field as key-value pairs after a ?:
mypage.html?customer-name=Alex&pickup-day=saturday&items=loafKeys come from the name attributes, values from what the visitor typed. This is also the proof that unnamed fields vanish - anything without a name simply isn’t in that URL.
| GET | POST | |
|---|---|---|
| Where the data goes | In the URL, visible & bookmarkable | Inside the request, out of the URL |
| Right for | Searches, filters - the data is a question | Orders, logins, messages - changes or private data |
Revision summary
Section titled “Revision summary”| Must-know | One-line recall |
|---|---|
| Three languages | HTML = structure, CSS = looks, JS = behavior - one job each |
| What HTML is | HyperText Markup Language - labels saying what content is, never how it looks |
| Document anatomy | doctype → html → head (about the page) → body (what’s seen) |
| The two meta lines | charset="UTF-8" (alphabet) + viewport (real width on phones) - fixed in every head |
| Element vs tag vs attribute | Element = markup + content; tag = the bracketed marker; attribute = name="value" in the opening tag |
| Void elements | <meta>, <img>, <br>, <input> - no content, no closing tag |
| Nesting rule | Opened last, closed first - boxes inside boxes, never overlapping |
| Headings | h1-h6 are a ranked outline, not sizes; one h1, don’t skip levels |
| Whitespace collapses | Runs of spaces/newlines become one space - only elements create structure |
| strong vs em | strong = important (survives out of context); em = vocal stress |
| Lists | ul (order-free), ol (order matters), li items; only li directly inside |
| Links | <a href>; relative inside the project, absolute only for external |
| Images | <img src alt> void; every image gets an alt sentence; lowercase names; compress |
| Semantic structure | header nav main section article footer = landmarks; div/span only when no meaning fits |
| Block vs inline | Block = own line, full width; inline = flows in text, content-width |
| index.html | Server returns a folder’s index.html by default - that’s the home page |
| Tables | table/thead/tbody/tr/th/td for data only, never layout |
| Form basics | form frames it; every input needs a wired label and a name (no name → sends nothing) |
| Input types | text, email, radio (one, shared name), checkbox (many), plus textarea/select |
| On submit | action = where, method = how; GET = data in URL, POST = data hidden; real validation is server-side |