Certification & Online Trace Collection · service active
WACZ · ISO 28500/ eIDAS timestamping/ Client area
C.E.R.T.O.
Sign in Register free
IT EN
C.E.R.T.O. / Modules / Web pages
01 · WEB

Web pages

Capture a web page exactly as it appears in that moment — with everything it is made of — and seal it before it can change or vanish.

Built-in forensic browser: target URL bar and a “Start acquisition” button. With “Screenshot” the operator captures the screen whenever they deem it appropriate while interacting with the page; navigation stays confined to the declared scope.

C.E.R.T.O.'s Web pages module performs the forensic acquisition of websites and web pages: it captures a page in the exact state it appears — text, images, code, resources and network traffic — and seals it as digital evidence before it can change or vanish. It is the tool designed for court-appointed and party experts, lawyers and law enforcement who need to freeze online content — a defamatory post, a copyright infringement, an e-commerce page, a review — with a full chain of custody and technical validity.

Key features

What this module does.

  • Replayable archive in WACZ format (ISO 28500 / WARC).
  • Built-in browser with HD video recording and system audio.
  • Full-page screenshots, HTTP traffic (HAR) and SSL/TLS analysis.
  • Multi-page with a capture-completeness engine.
  • BagIt 1.0 bundle signed with Ed25519, double RFC 3161 timestamp and CASE/UCO.
Final forensic report: the interactive index-interactive.html dashboard of the bundle — summary, pages, screenshots, network, certificates, completeness, hash inventory and integrity check.
Forensic pipeline

How the module operates.

A repeatable, documented procedure: from time synchronisation to cryptographic sealing, every step leaves a verifiable trace inside the bundle.

01 · NTP

Synchronised time

Multi-source NTP sync (Google/Cloudflare/pool) with documented offset and roundtrip: the capture window is anchored.

02 · SCOPE

Confined navigation

Built-in forensic browser with declared object and scope (domain, domain+subdomains or free navigation); every in-scope page becomes an exhibit.

03 · CAPTURE

Passive capture

Passive traffic capture via Chrome DevTools Protocol (Network domain, no request interception) — no page alteration.

04 · RENDER

Visual states & DOM

For each page: viewport screenshot, full-page screenshot (scroll-and-stitch) and a DOM snapshot after JavaScript execution.

05 · WACZ

Replayable archive

Self-contained WACZ/WARC packaging (ISO 28500): the navigation replays offline, byte for byte, with ReplayWeb.page.

06 · COMPLETE

Completeness

Session-aware completeness engine: compares seen vs captured resources, recovers the missing ones and declares the unrecoverable.

07 · NETWORK

Network & TLS

Complete W3C HAR, DNS, WHOIS, traceroute and X.509 certificate chains (leaf + intermediates + root) for every host contacted.

08 · HASH

Fingerprints

Quadruple MD5 + SHA-1 + SHA-256 + SHA-512 (FIPS 180-4) hash of every file, inventoried in file-hashes.json.

09 · SIGN

Signature & double timestamp

manifest.json signed with Ed25519 (RFC 8032) bound to the device + double RFC 3161 timestamp (inner anchor on the manifest, outer seal on the tag-manifest).

10 · SEAL

Sealed bundle

Everything is packaged into a BagIt 1.0 bundle (RFC 8493) with a CASE/UCO description and verify.sh / verify.bat verifiers.

Bundle contents

Everything that gets generated.

A single acquisition produces dozens of coordinated artefacts, each with a precise forensic role. They are organised into clearly-named folders inside data/.

Viewport screenshots

The visible screen during navigation (one frame per page), as JPEG watermarked with C.E.R.T.O., version, acquisition ID, URL and timestamp.

pages/NNN_…/screenshots/

Full-page screenshots

The whole page recomposed with scroll-and-stitch, beyond the fold, with the background flattened to white for colour fidelity. Generated post-hoc for each page.

pages/NNN_…/screenshots-fullpage/

DOM HTML snapshots

The DOM serialised after JavaScript execution (post-hydration), as actually rendered by the browser — not just the static source.

pages/NNN_…/html-snapshots/

WACZ archive

Web Archive Collection Zipped (ISO 28500): WARC + indexes, self-contained. Replays the whole navigation offline with ReplayWeb.page — it is the sealed media.

evidence/ReplayWebPage.wacz

Session video

Video recording (WebM) of the forensic browser during the acquisition, with system audio: the dynamic proof of what the operator saw and did.

evidence/video/

Network capture (HAR)

W3C HTTP Archive: complete record of requests/responses, real headers, timing and payload. Plus requests/responses/resources in JSON and statistics.

network/network.har

DNS · WHOIS · Traceroute

DNS resolution, WHOIS registry query (domain owner and registration data) and a map of the network hops from client to host.

network/dns-lookup.txt · whois.txt · traceroute.txt

SSL/TLS certificates

The full X.509 chains (leaf + intermediates + root, in PEM and human-readable) of every host contacted during the acquisition, with certificate details.

tls/certificates/

Cookies & JavaScript state

All active cookies with metadata (HttpOnly, Secure, SameSite, expiry) and the session storage / local storage snapshot at capture time.

network/cookies-detailed.json · evidence/javascript-state/

User interactions

Recording of clicks, scrolls, inputs and keystrokes during navigation (chain of custody of the operator's actions).

evidence/interactions/user-interactions.json

Resource map

The complete site-structure: every resource (CSS, JS, fonts, images) saved and organised by host, with the provenance metadata of each.

resources/site-structure/

Capture completeness

A report that honestly declares how many resources were seen, captured, recovered via session and how many remain unrecoverable, with the percentage.

network/completeness-report.json

Forensic report

The report in PDF and TXT (operator, scope, IP, NTP, SSL, inventory, forensic statements) with its own RFC 3161 timestamp (report.tsr).

reports/report.pdf · report.txt · report.tsr

Hash inventory

The list of every file with its quadruple of cryptographic hashes, the basis for an integrity check repeatable by anyone, even offline.

hashes/file-hashes.json

Social platforms

Comments are evidence. They must be taken where they live.

On a post, the defamation or the threat is often not in the content: it is in the comments below.

But the comments are not in the page. They appear little by little as you scroll, they hide behind a “show more replies”, and the platform keeps only a handful on screen at a time. A tool that photographs the screen collects what is visible at that moment: the rest does not exist, and it does not even show up as missing.

C.E.R.T.O. queries the same data source the platform uses to build the page and preserves every response exactly as received, with its provenance envelope: queried address, method, status, moment of receipt. Comment text, author, like count and date are not read off the screen — they come from the service, as the service declares them.

Replies do not come with the list: every comment that has them carries its own continuation, which must be followed separately. C.E.R.T.O. follows them, and when a list runs out it says so — with the reason why.

Instagram

Comments on posts and reels. Profile with the content grid, carousels included. Stories, isolating the selected item. Follower and following lists.

social/instagram/ · raw-api/

Facebook

Comments. Profile. Timeline posts. Friend list, with each person's identifier and picture.

social/facebook/ · raw-api/

TikTok

Comments on videos and photo posts. The video itself, downloaded as a standalone file with its own hash. Profile and published content.

social/tiktok/ · evidence/video-resources/

X (Twitter)

Posts with their replies, followed page after page to the end. Profile with published content, metrics and picture.

social/x/ · raw-api/

YouTube

Comments and replies on videos and shorts, distinguished by level. Channel with every thumbnail at full quality. The film downloaded separately.

youtube-comments/ · raw-api/

Messenger & Direct

Private Messenger and Instagram Direct conversations, extracted without having to scroll through them one by one.

social/messaging/

The raw response stays in the bundle

Every service response is preserved verbatim in raw-api/ with its provenance envelope. Anyone contesting the data can go back to the source, rather than trusting our table.

The document declares its limits

How many items the platform declares, how many were collected, through which channel, and why collection stopped. An incomplete collection is stated, not hidden.

Dates as the platform states them

Many platforms do not state an absolute date but a distance from now (“7 days ago”). The value is reported as received: converting it would be our interpretation, not a fact.

The bundle is browsable

Thousands of messages, and the one you need in three seconds.

Hundreds of comments and dozens of authors: a printed list is no use to anyone.

The interactive.html dashboard included in the bundle presents comments and replies as real tables, where you can search and sort. It works offline, on a double click, even years from now and on a computer without C.E.R.T.O. installed.

  • Search in the message text, in the author name and by identifier. The result count updates as you type.
  • Sorting on every column: author, kind, date, likes, text. Relative dates sort the right way round, most recent first.
  • Comments and replies kept distinct: you see at a glance what is a reply and which message it hangs from.
  • Comment screenshots: the gallery of frames taken during extraction, plus a single PDF carrying the hash of each image and of the PDF itself.
  • Media page: every image in the acquisition with MD5, SHA-1 and SHA-256, declared provenance and search by hash too.
  • Methodological caveats within reach but out of the way: they open from a marker, grouped by topic.
Technical honesty

What the archive replays, and what it does not.

On a traditional website the replay is faithful. On an application there is a limit, and it deserves an explanation.

The WACZ archive keeps resources under the address they were requested with, so the page reassembles as it was: structure, stylesheets, fonts, images. With X, Instagram, Facebook or TikTok things change — the limit is not ours, and no tool today gets past it.

Why an application does not replay

On replay the site's own code is re-executed: it redraws the page from scratch and asks the service for data that is not in the archive — because nobody had asked for it at that moment. The result is a page that complains, not a false page.

Hence the snapshot

Alongside the network capture, the archive holds the page as the operator saw it, with no executable code: it cannot rewrite itself because it has nothing to run. Next to the file sits a bilingual note explaining what each of the two is.

The film does not replay — so it is downloaded

Video streams travel under signed, expiring addresses, and the player demands a live answer from the service: no archive can replay them, not even by keeping the bytes. C.E.R.T.O. downloads the film separately, as a standalone file with its own hash.

None of this diminishes the evidence. The content sits in the bundle three times over, by independent routes: in the service's raw response, in the frames captured during extraction, and in the page archive. Each has its own hash, and the three corroborate one another. An archive that does not replay is not an empty archive: it holds what went over the wire at that moment, address by address.

Self-validation

A bundle that proves itself.

The bundle does not need C.E.R.T.O. to be validated: anyone, even years from now, can verify its authenticity with standard tools. The BagIt 1.0 structure and the interactive dashboard make it self-explanatory.

  • index-interactive.html — the navigable offline dashboard of the bundle: summary, visited pages, screenshots, video, network, certificates, completeness, hash inventory and client-side integrity check.
  • manifest.json signed with Ed25519 (RFC 8032), bound to the identity of the device registered at first launch.
  • Double RFC 3161 timestamp: inner anchor on data/tsa.tsr and outer seal on tagmanifest-sha256.txt.tsr. Free cascade Sectigo→DigiCert→GlobalSign; optional qualified eIDAS InfoCert.
  • manifest-sha256.txt and tagmanifest-sha256.txt (RFC 8493): fixity of the payload and of the control files; no file can be added or altered without the check failing.
  • metadata/evidence.case.jsonldCASE 1.3 / UCO 1.4 description of the evidence, and tsa-ca.pem for verifying the timestamp even offline.
  • verify.sh / verify.bat — standalone verifiers: they recompute the hashes, check the double timestamp and the signature, and declare “VALID BUNDLE”.
FAQ

Frequently asked questions

Forensic web page acquisition, evidence validity, the WACZ format and bundle verification: the most common questions.

What is forensic web page acquisition?
It is the capture of a web page in the exact state it appears at a given moment — content, code, resources, network traffic and certificates — cryptographically sealed so it can be used as digital evidence and verified by third parties.
What is the difference between a screenshot and a forensic acquisition?
A screenshot is just an image, easily manipulated and without context. C.E.R.T.O.'s forensic acquisition also collects the rendered HTML, the replayable WACZ archive, the HTTP traffic (HAR), the TLS certificates, cryptographic hashes and a double RFC 3161 timestamp, with a verifiable chain of custody.
Is the acquisition valid as evidence in court?
The bundle is produced according to recognised standards (ISO/IEC 27037, BagIt RFC 8493, RFC 3161, CASE/UCO) with an Ed25519 signature and a double timestamp. Its authenticity and integrity can be verified by anyone, even offline, making it suitable for expert and court use. The final assessment always rests with the adjudicating authority.
What is a WACZ archive?
WACZ (Web Archive Collection Zipped, based on WARC / ISO 28500) is a standard format that bundles the entire navigation — pages, resources and traffic — into a single self-contained file, replayable offline byte for byte with ReplayWeb.page.
How is the bundle's authenticity verified?
Every bundle contains verify.sh / verify.bat and an index-interactive.html dashboard: they recompute the hashes, check the Ed25519 signature and the double RFC 3161 timestamp and declare whether the bundle is valid. Verification needs neither C.E.R.T.O. nor an internet connection.
Does C.E.R.T.O. acquire the comments on a social post?
Yes, and it is often the part that matters: on Instagram, Facebook, TikTok, X and YouTube the comments — and the replies to comments — are collected by querying the same data source the platform uses to build the page, not by photographing the screen. Every service response stays in the bundle exactly as received, with the queried address and the moment of receipt. The document states how many items the platform announces, how many were collected and why collection stopped.
How do you find one message among hundreds of comments?
The interactive.html dashboard included in the bundle presents comments and replies in tables with search on text, author and identifier, and sorting on every column. It works offline, on a double click, even on a computer without C.E.R.T.O. installed. There is also the gallery of frames captured during extraction, with a PDF carrying their hashes, and a page listing every image in the acquisition, searchable by hash as well.
Why does the WACZ archive of a social platform not replay like the original site?
Because on replay the site's own code is re-executed: it redraws the page from scratch and asks the service for data that is not in the archive, since nobody had asked for it at that moment. It is a limitation of the format when facing applications, not of C.E.R.T.O., and no tool today gets past it. That is why the archive also holds a snapshot of the page as the operator saw it, with no executable code, and a bilingual note beside it. The content remains in the bundle by three independent routes anyway: the service's raw response, the screenshots and the archive itself.
Is the video of a post acquired?
Yes, but not inside the page archive: video streams travel under signed, expiring addresses and no archive can replay them. C.E.R.T.O. downloads the film separately, as a standalone file with its own cryptographic hash, so it stays playable with any player even years from now.
What content is acquired?
Viewport and full-page screenshots, a DOM snapshot after JavaScript execution, the WACZ archive, the session video, HAR network traffic, DNS/WHOIS/traceroute, SSL/TLS certificates, cookies and JavaScript state, user interactions, the resource map and a PDF forensic report, with a complete hash inventory.

Collect evidence with the Web pages module.

Register for free and download C.E.R.T.O. Desktop for Windows and macOS from your client area.