Certification & Online Trace Collection · service active
WACZ · ISO 28500/ eIDAS timestamping/ Client area
C.E.R.T.O.
Sign in Register free
IT EN
C.E.R.T.O. / Modules / Web pages / X (Twitter)
PLATFORM · X

X (Twitter)

Posts with their replies, followed page after page to the end. Profile with published content, metrics and picture.

It switches itself on
  1. You paste the address into the forensic browser.
  2. You start a web page acquisition, as always.
  3. C.E.R.T.O. recognises the platform and does the rest.

No option to tick, no separate module to buy, no extra cost: it is the Web pages module, at its own rate.

On X what matters is almost always the chain: the post and the replies piling up beneath it, often from different authors at different times. The replies are not in the page: they are requested from the service one page at a time, following a continuation.

There is a trap measured in the field: X keeps offering a continuation even when it has nothing left to give. A loop that stops only at the missing continuation keeps asking in vain until the network gives way, and then writes in the bundle that collection failed — when in fact it had finished long before. C.E.R.T.O. stops after three pages with no novelty and states which was the last one with content.

What goes into the bundle

What gets acquired

Only what the program actually does: every item matches a collection that exists in the code and has been tested on real acquisitions.

Post and replies

Content, author, metrics and the replies followed to the end, with the reason for stopping stated.

The post video

Downloaded separately as a standalone file with its own hash, when the post carries one.

Profile

Details and published content. The details are taken from the author node of the content, and the bundle states that this is the source.

No duplicate frames

Replies arrive from the service without being drawn on screen: a periodic capture would photograph the same thing over and over. Only the frame that changes is kept, and how many were removed is written in the bundle.

Recognised addresses

What you can paste

Recognition looks at the domain and the shape of the address. If it recognises nothing, it stays an ordinary web page acquisition: nothing is lost.

  • x.com/<utente>/status/<id>post with its replies
  • x.com/<utente>profile and published content
  • twitter.com/…the legacy domain still works
In the bundle

Where what it collects ends up

Folders inside the BagIt bundle, alongside everything else from the web acquisition: WACZ archive, screenshots, network traffic, certificates, report.

  • social/x/
  • raw-api/

The bundle's interactive.html dashboard presents them in tables with search and sorting, and works offline, on a double click, even on a computer without C.E.R.T.O. installed.

The raw response stays in the bundle

Where collection goes through a service, every response is preserved verbatim with its provenance envelope: queried address, method, status, moment of receipt. Anyone contesting the data can go back to the source rather than trusting our table.

Sealed like every acquisition

A BagIt 1.0 bundle with an Ed25519 signature, a double RFC 3161 timestamp and a CASE/UCO description, verifiable by anyone, offline, with the included verifiers.

It says when it stops

The document reports how many items the platform announces, how many were collected and why collection stopped. A stated limit is worth more than an asserted completeness.

FAQ

Frequently asked questions

The questions we are asked most often about this platform.

Are all the replies to a post acquired?
Replies are followed page after page as long as the service returns new ones. When it stops, C.E.R.T.O. stops too and states in the bundle how many pages it requested, which was the last one with content and why it stopped. It never writes “all acquired” without having verified it.
Does it work on x.com or twitter.com?
On both, including the www and mobile forms: an address copied from a phone is recognised just like one copied from a computer.
Why does the WACZ archive of X not replay like the site?
Because on replay the site's own code is re-executed, redraws the page from scratch and asks for data that is not in the archive, since nobody had asked for it at that moment. It is a limitation of the format when facing applications, not of C.E.R.T.O., and no tool today gets past it. That is why the content stays in the bundle by three independent routes: the raw service response, the screenshots and a snapshot of the page with no executable code.
Same module

The other platforms

All inside the Web pages module, all recognised from the address, all at the same rate.

Collect evidence from X (Twitter).

Register for free and download C.E.R.T.O. Desktop for Windows and macOS from your client area.