Showing posts with label Forensics. Show all posts
Showing posts with label Forensics. Show all posts

Friday, January 2, 2015

Forensic Analysis of Microsoft office OOXML/OpenXML files using Python


Microsoft started supporting Office Open XML format from Microsoft Office 2007 release onwards.
Office Open XML is also know as OOXML or OpenXML. All office files, say, docx, xlsx, pptx are zipped content of XML files.

In this blog we will try to read find xml files part of the office file archive and look at directory structure. We will parse office files using Python.

Libraries used
        import zipfile, sys
        import libxml2, datetime
zipfile       for parsing zipfile
sys             to use exit() function
libxml2     for parsing XML files
datetime    for printing date/time

Read office file(which is zipped indeed)
file = zipfile.ZipFile(sys.argv[1],"r")
Where sys.argv[1] is the command line argument we pass to the program as shpown below
./office.py file.docx

Loop through all the files and print
        for name in file.namelist():
            print "file name:" + name
To read XML file
        xmlbuf = file.read(name)
file.open() will not work

Read the XML file, we can also use  libxml2.parseMemory(xmlbuf)

        try:
            xmlf = libxml2.parseDoc(xmlbuf)
        except (libxml2.parserError, TypeError):
            print "Error loading core.xml"
            sys.exit()

Get the root element of the XML
        root = xmlf.getRootElement()
coreProperties in the case of core.xml

Call below function with root tag/element as argument
        recursive_find(root)
This is a recursive function which recursively gets all the XML tags.

Entity: line 1: parser error : Start tag expected, '<' not found
docProps/core.xml

To ignore above error use
        def noerr(ctx, str):
            pass
        libxml2.registerErrorHandler(noerr, None)
The error above is the major blocking point.

core.xml file part of the office file archive has details like file creators name, access time, modified time, last modified user's name etc. We will try to extract those values using Python. Apart from core.xml file we also print names, file sizes, time stamps of the files part of the archive.

Final result will look like


http://en.wikipedia.org/wiki/Office_Open_XML
http://www.ecma-international.org/publications/standards/Ecma-376.htm
http://msdn.microsoft.com/en-us/library/dd908153(v=office.12).aspx

Saturday, July 26, 2014

Incidence Response: Important Linux Commands and Log Files

Most of the log files are located at
/var/log/

btmp, utmp, wtmp
last -f /var/log/btmp | more
last
recent login information for all the users
lastlog                

/var/log/secure       contains information about authentication and authorization

auth.log
maillog

Saturday, October 2, 2010

Forensics 2: Identifying File System and Extracting it

The advantages of analyzing disk images are that the investigators can:
a) preserve the digital crime-scene
b) obtain the information in slack space
c) access unallocated space, free space, and used space
d) recover file fragments, hidden or deleted files and directories
e) view the partition structure and
f) get date-stamp and ownership of files and folders.

Here we will try to concentrate on extracting the File System if any from the image for analysis available from the Crime Scene.

Lets check the md5 hash of the image under analysis for integrity purposes. The md5 hash algorithm produces a 128 bit “fingerprint” of a file, also known as a message digest. To view the md5 hash value assigned to a given file, the md5sum utility can be used.Lets check the file type of the image under analysis by using file command. The file command works by testing “arguments” within a file, and will then classify the file as whichever file type the file command sees fit. We see from the output of the file command that the image file contains an x86 boot sector. The boot sector of a computer is a primary starting point for an OS. The operating system will start at the boot loader, and the machine will read the first 512 bytes of the disk, which is known as the boot sector. The first 512 Bytes (boot sector) will be loaded into memory and will then be executed. This will initiate the boot process.

The x86 boot sector type message was obtained because the magic number 0xAA55 value is located at the 0x1FE offset within the image; defined in the file “/usr/share/file/magic” which is used by file command.

Determining the File System type of the Image
Lets run mmls utility to determine the File System type of the given image extracted by using dd command as shown below.-t Specify the media management type (dos, mac, bsd etc)

We see above that the NTFS (New Technology File System) partition begins at sector 63 (to see this look at the last column in the row where it says NTFS (0x07). Now look to the left in the start column of the row NTFS and we can see the value 0000000063). So for all the Sleuth Kit commands we need to specify an offset of 63 if the file used is raw image.
MMLS is a forensics utility that query’s an image file, and prints the partition tables and disk labels. This command is very useful when attempting to determine at which sector a partition begins and ends. We see that there is a NTFS file system on this image. We use the –t dos switch to tell mmls to utilize a dos based architecture for the file system.


File system is extracted using dd.exe command. Input file is the raw image collected from the machine which is under forensic investigation. Block size used to extract File system is 512 bytes and skipped 62 sectors because our NTFS File System is starting after those sectors.

Thus extracted File System image can be mounted by using mount command, we can check the mounted File System using fdisk -l command.




After extracting the image calculate md5 of the extracted NTFS File System image for integrity purposes.



Extracting the File System from the image


-b partition sizes in bytes
-r Recurse into DOS partitions and look for other partition tables.
-v verbose

Tuesday, September 14, 2010

Forensics 1: Extracting an Image for Investigation

Forensic investigations are usually performed on Static Data (images). Many open source (TSK) and commercial tools (Encase) are available for forensic analysis of a given image.
Lets look at how to take the image of a drive, hard disk, partition etc. Few tools which can be used are dd, windd etc.
Well, what is an image? Image is a bit-by-bit copy of the Hard Disk.
I used dd.exe command for taking the image of the computer under investigation.
dd command is found by default in Linux. On windows we can obtain the binary from The Sleuth Kit (TSK) or comes by default if Cygwin is installed.
First, lets list all the available drives (A:, B:, C: etc.,) or partitions on the machine where we want to collect image.
Below is the snapshot of the dd command used for extracting the image for investigation.
We should be very cautious while collecting the image for investigation because nothing should be changed on the machine under analysis. So most of the time we should use CD with all the tools and redirect the image to external drive or network share for saving rather than saving the image on local machine.
dd command can also be used to extract a File System from Raw image.
This is just a high level overview of Forensics will come up with more articles.
For further reading you can start from
http://en.wikipedia.org/wiki/Computer_forensics