So far, my thesis is going well. The most difficult part is holding myself back from beginning data processing. I’ve scoured the internet and found that, as I suspected, there is little to no research on standard data management for internally derived datasets in cultural heritage. This works in my favor, as I can use this project to fill that gap.
My plan is to take GIS data ethics from archaeology and data management strategies from machine learning, weave them together, and create a framework to serve the needs of cultural heritage institutions. On that note, I need to define a word for cultural heritage that excludes archaeology. I find that there is a conflation between archaeological GIS and historical GIS. Actually, I would go so far as to say there is a difference between historical GIS and the point of this project. In archaeological GIS the spatial data comes from the object's location. In historical GIS there is both intrinsic and extrinsic spatial data. This generally takes one of two forms: A map (intrinsic) or provenance (extrinsic). What this project works with is a collection with intrinsic textual spatial data. That is to say there is no spatial data within the object or about the object, we are interpreting the text to obtain spatial data. Technically, this is intrinsic data, however it is clearly different from a map.
Returning to the point, I have been messing around with the 2002 WWII enlistment data. I converted the ASCII code to hexadecimal, which with the help of Claude was quite easy, but did increase the file size to 5 gb (yes gb). I was expecting 50-60 gb. Now that the data is in hex, it needs to be decoded. To do so I will need to (a) figure out a way to extract all the codes from the documentation quickly and (b) how to effectively chunk the data to avoid sitting for hours while the code runs. Outside of the lit review the next step is figuring out OCR.
No comments:
Post a Comment