C.E.R.T.O. — AVVERTENZA SULLA RIPRODUZIONE DELL'ARCHIVIO WEB (WACZ)
====================================================================

Il file ReplayWebPage.wacz e un archivio web (WACZ, contenente WARC ISO 28500).
Si apre con un lettore di archivi, per esempio https://replayweb.page

CHE COSA ASPETTARSI

Se la pagina acquisita era un sito ordinario, si riproduce normalmente.

Se invece era un'APPLICAZIONE WEB (X/Twitter, Instagram, Facebook, TikTok,
Google Maps e simili), la pagina originale puo apparire per un istante e poi
mostrare un errore della piattaforma, del tipo "Something went wrong."

ELENCO DELLE PAGINE: SI OFFRE SOLO CIO CHE SI RIPRODUCE

Su alcune acquisizioni l'elenco delle pagine e stato RIPULITO: vi compaiono solo
le pagine che, aperte, mostrano davvero qualcosa. Le pagine vive di
un'applicazione, che in riproduzione restano vuote, e gli snapshot duplicati
della stessa scheda NON sono elencati. Non sono pero stati tolti: i loro record
restano nell'archivio, con le loro impronte, e sono raggiungibili cercandone
l'indirizzo nella scheda "Resources" del lettore. Elencarli avrebbe promesso al
lettore una porta che non porta da nessuna parte; toglierli avrebbe sottratto
materiale al reperto. Non si e fatta ne l'una ne l'altra cosa.

PERCHE ACCADE — E PERCHE NON E UN DIFETTO DELL'ACQUISIZIONE

Quei siti non sono documenti: sono programmi. Alla riproduzione il loro codice
viene rieseguito, ridisegna la pagina da zero e chiede alla rete decine di
chiamate interne. L'archivio ne contiene solo quelle realmente emesse durante
l'acquisizione; le altre non possono esserci, perche in riproduzione la
piattaforma non e raggiungibile e nessuno puo rispondere. Il programma, non
ricevendo risposta, mostra il proprio errore e cancella cio che aveva disegnato.

Nessun archivio di questi siti, prodotto con QUALSIASI strumento, si riproduce
per intero. E un limite della tecnologia di archiviazione applicata alle
applicazioni web, non una lacuna della raccolta.

DOVE GUARDARE INVECE

1) Nell'elenco delle pagine del lettore c'e una seconda voce, il cui titolo
   comincia con "[CERTO SNAPSHOT]" e il cui indirizzo termina con
   ".certo-snapshot-<data>.html". Quella e la pagina COME L'HA VISTA
   L'OPERATORE. ATTENZIONE: i siti che mostrano flussi lunghi collocano le
   righe al loro spostamento di scorrimento, quindi lo snapshot puo aprirsi
   AL DI SOPRA del contenuto: se la parte alta appare vuota, SCORRERE la
   pagina verso il basso. Il posizionamento non e stato alterato proprio per
   conservare la resa fedele.
   L'OPERATORE: e la stessa cattura, privata del codice che la ridisegnerebbe.
   Testi, immagini e stili sono quelli raccolti; nulla e stato aggiunto.

2) Il rapporto interattivo (interactive.html) riporta il contenuto in forma
   consultabile, con le impronte di ogni file.

3) Le risposte del servizio, quando il modulo di acquisizione le ha raccolte,
   sono conservate integre nella cartella raw-api/ con la busta di provenienza.

La cattura originale resta nell'archivio, invariata: serve a verificare che il
snapshot non sia stato alterato. Le uniche differenze fra i due sono
dichiarate nel sorgente dello snapshot stesso.

L'INDIRIZZO DELLE PAGINE DENTRO L'ARCHIVIO

Se l'indirizzo acquisito conteneva un frammento (la parte che segue il primo
"#", che alcuni siti aggiungono al «copia link»), dentro l'archivio le pagine
sono indirizzate SENZA di esso. Non e una modifica del reperto: il frammento
non viene mai spedito al server — resta nel browser — quindi non fa parte
dell'indirizzo con cui una risorsa puo essere richiesta, ne ritrovata in un
archivio. L'indirizzo verbatim indicato dall'operatore, frammento compreso,
resta invariato nel rapporto, nel manifest.json e in bag-info.txt (Source-URL).


C.E.R.T.O. — NOTICE ON WEB ARCHIVE (WACZ) REPLAY
====================================================================

The file ReplayWebPage.wacz is a web archive (WACZ, containing ISO 28500 WARC).
Open it with an archive replay tool, for example https://replayweb.page

WHAT TO EXPECT

If the captured page was an ordinary website, it replays normally.

If it was a WEB APPLICATION (X/Twitter, Instagram, Facebook, TikTok and the
like), the original page may appear for a moment and then show a platform error
such as "Something went wrong. Try reloading."

WHY THIS HAPPENS — AND WHY IT IS NOT AN ACQUISITION DEFECT

Such sites are not documents: they are programs. On replay their code runs
again, redraws the page from scratch and issues dozens of internal network
calls. The archive contains only those actually issued during the acquisition;
the others cannot be there, because on replay the platform is unreachable and
nobody can answer. Receiving no answer, the program shows its own error and
discards what it had drawn.

No archive of these sites, produced with ANY tool, replays in full. This is a
limitation of archiving technology applied to web applications, not a gap in
the collection.

WHERE TO LOOK INSTEAD

1) The replay tool's page list contains a second entry, titled starting with
   "[CERTO SNAPSHOT]" and with an address ending in
   ".certo-snapshot-<date>.html". That is the page AS THE OPERATOR SAW
   IT. NOTE: sites showing long feeds place rows at their scroll offset, so
   the snapshot may open ABOVE the content: if the top looks empty, SCROLL
   down. Positioning was deliberately left untouched to keep the rendering
   faithful.
   IT: the same capture, stripped of the code that would redraw it. Text,
   images and styles are the ones collected; nothing has been added.

2) The interactive report (interactive.html) presents the content in browsable
   form, with the hashes of every file.

3) Service responses, where the acquisition module collected them, are kept
   verbatim in the raw-api/ folder together with their provenance envelope.

The original capture remains in the archive, unchanged: it allows verifying
that the rendered document was not altered. The only differences between the
two are declared in the source of the rendered document itself.

HOW PAGES ARE ADDRESSED INSIDE THE ARCHIVE

If the acquired address carried a fragment (whatever follows the first "#",
which some sites append to their "copy link"), pages are addressed inside the
archive WITHOUT it. This is not an alteration of the evidence: a fragment is
never sent to the server — it stays in the browser — so it is not part of the
address by which a resource can be requested, nor looked up in an archive. The
verbatim address entered by the operator, fragment included, is left unchanged
in the report, in manifest.json and in bag-info.txt (Source-URL).
