diff options
| author | Rusty Wagner <rusty.wagner@gmail.com> | 2024-03-05 19:50:13 -0500 |
|---|---|---|
| committer | Rusty Wagner <rusty.wagner@gmail.com> | 2024-03-05 20:34:34 -0500 |
| commit | e093c21ed880ac3eb72119be15093ee04f8ce299 (patch) | |
| tree | 9f720ebdc0ae415734b1199ed341668c69710a94 /arch/armv7/thumb2_disasm/README.md | |
| parent | 0609276712622908254065546102381466033141 (diff) | |
Move architecture modules into the API repo
Diffstat (limited to 'arch/armv7/thumb2_disasm/README.md')
| -rw-r--r-- | arch/armv7/thumb2_disasm/README.md | 36 |
1 files changed, 36 insertions, 0 deletions
diff --git a/arch/armv7/thumb2_disasm/README.md b/arch/armv7/thumb2_disasm/README.md new file mode 100644 index 00000000..6d849e10 --- /dev/null +++ b/arch/armv7/thumb2_disasm/README.md @@ -0,0 +1,36 @@ +# ARM Thumb Decomposer/Disassembler +This is a disassembler for ARM Thumb (mixed 16-bit and 32-bit). Currently, it's scope does not contain Thumb2 or ThumbEE. + +# Terms +I'm using "Decomposer" to mean that instruction data is analyzed and a useful description of that instruction data is produced. +More specific, an "instruction info" struct is created, capturing information about the instruction like its source registers and such. +I'm using "Disassembler" to mean that this "instruction info" struct from the decompose stage can be parsed to generated a human readable string that we commonly associate with disassembly. +This contains the instruction mnemonic and operands and any annotations (like the S suffix or condition flag). + +# High Level Strategy +Capture as much as possible from the specification (ARM Architecture Reference Manual, ARMv7-A and ARMv7-R edition). +Currently that's being done in spec.graph. +Then, parse that information (see process.py) into C source (see generated.c). +Finally, add another source file that calls into generated.c to interface with the rest of Binary Ninja (haven't look at this yet). + +# Lower Level Strategy +The tables of instructions become nodes in a graph. +When one table refers to another, that's an edge to another table. +And when a table holds only references to instruction encodings, it's a terminal node. +So decomposing/disassembling is traversing the graph from root to tip. +The intermediate language used to capture this table/node info kind of gets into a tradeoff game. +On one hand, I want to be able to copy/paste as much as possible from the spec. +On the other, I want the language to be simple enough that I don't need to recall anything from CompSci to write a simple parser for it. +The parser can be written in a nice easy language too; here, python. + +# How To Actually Generate? +Just run generator.py. It will read spec.txt and write spec.cpp. + +# Notes +- '.n' and '.w' qualifier/specifier select the narrow and wide encodings +- 's' suffix on instructions means it updates the flags +- cmp,cmn,tst,teq are result-less forms of subs,adds,ands,eors, but only update flags (but don't require the extra 's' suffix) +- there are 4 flags N,Z,C,V for negative,zero,carry,overflow +- the 'c' on b<c> is a conditional execution code (14 total) that test the flags + - code can be 'eq', 'ne', 'cs'/'hs', 'cc'/'lo', 'mi', 'pl', 'vs', 'vc', 'hi', 'ls', 'ge', 'lt', 'gt', 'le', 'al'/'' + |
