Skip to content

Added fiwalk disk image demo - #40

Merged
ajnelson-nist merged 4 commits into
dfxml-working-group:mainfrom
bitsgalore:main
Nov 17, 2022
Merged

ajnelson-nist merged 4 commits into
dfxml-working-group:mainfrom
bitsgalore:main

Conversation

@bitsgalore

Copy link
Copy Markdown
Contributor

Added demo that shows how to use fiwalk on a disk image and report results to an output file

@simsong simsong left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Buffers the entire XML image in memory. It would be better to just have fiwalk dump to a file.

Comment on lines +9 to +15
with open(imageFile, "rb") as ifs:
fwOutBuffer = fiwalk.fiwalk_xml_stream(imagefile=ifs)
fwOut = fwOutBuffer.read()

# Write dfxml to output file
with io.open(outFile, "wb") as fOut:
fOut.write(fwOut)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the contribution. My concern about this is that it requires the entire XML stream to be buffered in RAM at once (in fwOut) but there may be no realistic way to do this. It will fail for large images, however, unless run on systems with a huge amount of available memory. (The XML format is not very efficient.)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If there's a more efficient way to do this I'm all for it, because yes I can see how this could easily go wrong for large images. I've never worked with io.BufferedReader objects before though, and I'm not quite sure how to do this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think this changes the status quo for how DFXML has been implemented to date. The read-write library dfxml/objects.py does some full-in-memory string construction as well.

My litmus test (not to include in CI!) would probably be seeing if this can run against the NPS 2TB image.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It’s complicated. I believe there is an iterator in the dfxml package that runs fiwalk as a sub process , parses it with Expat, and calls a provided call back with an object as the argument. This is the interface that all of my code used. But that was 10 years ago.

@simsong

simsong commented Nov 9, 2022

Copy link
Copy Markdown

@ajnelson - What's your take? Do we want this, or do we want a more efficient example?

@ajnelson-nist

Copy link
Copy Markdown
Contributor

I'm having a look at this now. It looks like mypy no longer likes a test library import strategy, so I'll be fixing that separately.

@ajnelson-nist

Copy link
Copy Markdown
Contributor

This branch contains two things:

  1. A catch-up merge with main so CI will pass again, after merging PR 41.
  2. Type annotations which I think answer the part of initial question from Issue 39 that bites me every few years: roughly "What is supposed to be the type in or out of this DFXML outermost method?" One of my type-review sprees was because after a decade I'd still managed to catch myself mixing up text vs. binary IO in one of the XML ElementTree methods in objects.py.

Now, reviewing the other demo apps, there are other memory-efficient demonstrations that use the read-only SAX model. See e.g. demo_plot_times.py. The tradeoff with those is you have to do your analysis with a SAX mentality, i.e. with XML event-triggered functions.

The main contribution of demo_fiwalk_diskimage.py is that it opens Fiwalk in a subprocess, holds the XML as a binary stream in memory, and then dumps the stream to disk. So, it's somewhat like a pipe for command-line Fiwalk through Python. It doesn't demonstrate manipulating or parsing the XML stream with e.g. the Python XML modules, though each should be able to start from that stream.

I think the demo's worth including within DFXML because there doesn't seem to be a demonstration of fiwalk.fiwalk_xml_stream, but the demo needs comments inlined that summarize that last paragraph.

If you agree, @bitsgalore , would you mind merging the add_fiwalk_disk_image_demo branch into your main and adding a comment or two to the demo script?

@bitsgalore

Copy link
Copy Markdown
Contributor Author

@ajnelson-nist With some delay I just added some comments, let me know if this is what you had in mind.

As an aside, for my own use case, in the end I decided to wrap the fiwalk tool directly in a custom wrapper using the -X option. This way I now write fiwalk's output directly to a file, without using any in-memory Python buffers. This also means I don't really need dfxml_python for my particular case, so I'm not entirely sure how useful the demo is in the end.

Anyway, I'll leave it up to your whether you want to proceed with merging this PR or not.

@ajnelson-nist

Copy link
Copy Markdown
Contributor

Thank you for the summary. I will merge this once CI finishes. It's good to have the extra annotated demo, and I personally favor more type-annotating.

@ajnelson-nist
ajnelson-nist merged commit bf49107 into dfxml-working-group:main Nov 17, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants