4 ms·
The Internet Archive or the Library of Congress may archive much of the material: http://www.readwriteweb.com/archives/twitters_entire_archive_headed_to_the_li
by RobGR 16y ago
The Internet Archive or the Library of Congress may archive much of the material:
http://www.readwriteweb.com/archives/twitters_entire_archive_headed_to_the_library_of_c.php http://www.readwriteweb.com/archives/twitters_entire_archive...
http://www.archive.org/ http://www.archive.org/
If you want to archive it in a searchable way yourself, you may find that a very technically daunting task. People use large clusters of computers and pretty advanced technology to attempt stuff like this.
On the other hand, if you restrict the amount of data you want to ingest, you can get pretty far just by configuring existing software. You could set up Drupal on a hosting service that include Apache Solr search/indexing, such as Acquia or Pantheon, and then configure Drupal to pull in information from selected twitter searches and feeds of various blogs, and archive it.
Eventually you would run into space and volume restrictions. If you picked enough data sources, you eventually might max out the rate of wrties to MySQL. But you can go pretty far just by point and click configuring what exists.