64tass v1.59 r3120 reference manual

This is the manual for 64tass, the multi pass optimizing macro assembler for the 65xx series of processors. Key features:

Contrary how the length of this document suggests 64tass can be used with just basic 6502 assembly knowledge in simple ways like any other assembler. If some advanced functionality is needed then this document can serve as a reference.

This is a development version. Features or syntax may change as a result of corrections in non-backwards compatible ways in some rare cases. It's difficult to get everything right first time.

Project page: https://sourceforge.net/projects/tass64/

The page hosts the latest and older versions with sources and a bug and a feature request tracker.



Usage tips

64tass is a command line assembler, the source can be written in any text editor. As a minimum the source filename must be given on the command line. The -a command line option is highly recommended if the source is Unicode or ASCII.

64tass -a src.asm

There are also some useful parameters which are described later.

For comfortable compiling I use such Makefiles (for make):

demo.prg: source.asm macros.asm pic.drp music.bin
        64tass -C -a -B -i source.asm -o demo.tmp
        pucrunch -ffast -x 2048 demo.tmp >demo.prg

This way demo.prg is recreated by compiling source.asm whenever source.asm, macros.asm, pic.drp or music.bin had changed.

Of course it's not much harder to create something similar for win32 (make.bat), however this will always compile and compress:

64tass.exe -C -a -B -i source.asm -o demo.tmp
pucrunch.exe -ffast -x 2048 demo.tmp >demo.prg

Here's a slightly more advanced Makefile example with default action as testing in VICE, clean target for removal of temporary files and compressing using an intermediate temporary file:

all: demo.prg
        x64 -autostartprgmode 1 -autostart-warp +truedrive +cart $<

demo.prg: demo.tmp
        pucrunch -ffast -x 2048 $< >$@

demo.tmp: source.asm macros.asm pic.drp music.bin
        64tass -C -a -B -i $< -o $@

.INTERMEDIATE: demo.tmp
.PHONY: all clean
clean:
        $(RM) demo.prg demo.tmp

It's useful to add a basic header to your source files like the one below, so that the resulting file is directly runnable without additional compression:

*       = $0801
        .word (+), 2005  ;pointer, line number
        .null $9e, format("%4d", start);will be sys 4096
+       .word 0          ;basic line end

*       = $1000

start   rts

A frequently coming up question is, how to automatically allocate memory, without hacks like *=*+1? Sure there's .byte and friends for variables with initial values but what about zero page, or RAM outside of program area? The solution is to not use an initial value by using ? or not giving a fill byte value to .fill.

*       = $02
p1      .addr ?         ;a zero page pointer
temp    .fill 10        ;a 10 byte temporary area

Space allocated this way is not saved in the output as there's no data to save at those addresses.

What about some code running on zero page for speed? It needs to be relocated, and the length must be known to copy it there. Here's an example:

        ldx #size(zpcode)-1;calculate length
-       lda zpcode,x
        sta wrbyte,x
        dex             ;install to zero page
        bpl -
        jsr wrbyte
        rts
;code continues here but is compiled to run from $02
zpcode  .logical $02
wrbyte  sta $ffff       ;quick byte writer at $02
        inc wrbyte+1
        bne +
        inc wrbyte+2
+       rts
        .endlogical

The assembler supports lists and tuples, which does not seems interesting at first as it sound like something which is only useful when heavy scripting is involved. But as normal arithmetic operations also apply on all their elements at once, this could spare quite some typing and repetition.

Let's take a simple example of a low/high byte jump table of return addresses, this usually involves some unnecessary copy/pasting to create a pair of tables with constructs like >(label-1).

jumpcmd lda hibytes,x   ; selected routine in X register
        pha
        lda lobytes,x   ; push address to stack
        pha
        rts             ; jump, rts will increase pc by one!
; Build a list of jump addresses minus 1
_       := (cmd_p, cmd_c, cmd_m, cmd_s, cmd_r, cmd_l, cmd_e)-1
lobytes .byte <_        ; low bytes of jump addresses
hibytes .byte >_        ; high bytes

There are some other tips below in the descriptions.


Expressions and data types

Integer constants

Integer constants can be entered as decimal digits of arbitrary length. An underscore can be used between digits as a separator for better readability of long numbers. The following operations are accepted:

Integer operators and functions
x + yadd x to y2 + 2 is 4
x - ysubtract y from x4 - 1 is 3
x * ymultiply x with y2 * 3 is 6
x / yinteger divide x by y7 / 2 is 3
x % yinteger modulo of x divided by y5 % 2 is 1
x ** yx raised to power of y2 ** 4 is 16
-xnegated value-2 is -2
+xunchanged+2 is 2
~x-x - 1~3 is -4
x | ybitwise or2 | 6 is 6
x ^ ybitwise xor2 ^ 6 is 4
x & ybitwise and2 & 6 is 2
x << ylogical shift left1 << 3 is 8
x >> yarithmetic shift right-8 >> 3 is -1

Integers are automatically promoted to floats as necessary in expressions. Other types can be converted to integer using the integer type int.

Integer division is a floor division (rounding down) so 7 / 4 is 1 and not 1.75. If ceiling division is required (rounding up) that can be done by negating both the divident and the result. Typically it's done like 0 - -5 / 4 which results in 2.

        .byte 23        ; as unsigned
        .char -23       ; as signed

; using negative integers as immediate values
        ldx #-3         ; works as '#-' is signed immediate
num     = -3
        ldx #+num       ; needs explicit '#+' for signed 8 bits

        lda #((bitmap >> 10) & $0f) | ((screen >> 6) & $f0)
        sta $d018

Bit string constants

Bit string constants can be entered in hexadecimal form with a leading dollar sign or in binary with a leading percent sign. An underscore can be used between digits as a separator for better readability of long numbers. The following operations are accepted:

Bit string operators and functions
~xinvert bits~%101 is ~%101
y .. xconcatenate bits$a .. $b is $ab
y x nrepeat%101 x 3 is %101101101
x[n]extract bit(s)$a[1] is %1
x[s]slice bits$1234[4:8] is $3
x | ybitwise or~$2 | $6 is ~$0
x ^ ybitwise xor~$2 ^ $6 is ~$4
x & ybitwise and~$2 & $6 is $4
x << ybitwise shift left$0f << 4 is $0f0
x >> ybitwise shift right~$f4 >> 4 is ~$f

Length of bit string constants are defined in bits and is calculated from the number of bit digits used including leading zeros.

Bit strings are automatically promoted to integer or floating point as necessary in expressions. The higher bits are extended with zeros or ones as needed.

Bit strings support indexing and slicing. This is explained in detail in section Slicing and indexing.

Other types can be converted to bit string using the bit string type bits.

        .byte $33       ; 8 bits in hexadecimal
        .byte %00011111 ; 8 bits in binary
        .text $1234     ; $34, $12 (little endian)

        lda $01
        and #~$07       ; 8 bits even after inversion
        ora #$05
        sta $01

        lda $d015
        and #~%00100000 ;clear a bit
        sta $d015

Floating point constants

Floating point constants have a radix point in them and optionally an exponent. A decimal exponent is e while a binary one is p. An underscore can be used between digits as a separator for better readability. The following operations can be used:

Floating point operators and functions
x + yadd x to y2.2 + 2.2 is 4.4
x - ysubtract y from x4.1 - 1.1 is 3.0
x * ymultiply x with y1.5 * 3 is 4.5
x / yinteger divide x by y7.0 / 2.0 is 3.5
x % yinteger modulo of x divided by y5.0 % 2.0 is 1.0
x ** yx raised to power of y2.0 ** -1 is 0.5
-xnegated value-2.0 is -2.0
+xunchanged+2.0 is 2.0
~xalmost -x~2.1 is almost -2.1
x | ybitwise or2.5 | 6.5 is 6.5
x ^ ybitwise xor2.5 ^ 6.5 is 4.0
x & ybitwise and2.5 & 6.5 is 2.5
x << ylogical shift left1.0 << 3.0 is 8.0
x >> yarithmetic shift right-8.0 >> 4 is -0.5

As usual comparing floating point numbers for (non) equality is a bad idea due to rounding errors.

The only predefined constant is pi.

Floating point numbers are automatically truncated to integer as necessary. Other types can be converted to floating point by using the type float.

Fixed point conversion can be done by using the shift operators. For example an 8.16 fixed point number can be calculated as (3.14 << 16) & $ffffff. The binary operators operate like if the floating point number would be a fixed point one. This is the reason for the strange definition of inversion.

        .byte 3.66e1       ; 36.6, truncated to 36
        .byte $1.8p4       ; 4:4 fixed point number (1.5)
        .sint 12.2p8       ; 8:8 fixed point number (12.2)

Character string constants

Character strings are enclosed in single or double quotes and can hold any Unicode character.

Operations like indexing or slicing are always done on the original representation. The current encoding is only applied when it's used in expressions as numeric constants or in context of text data directives.

Doubling the quotes inside string literals escapes them and results in a single quote.

Character string operators and functions
y .. xconcatenate strings"a" .. "b" is "ab"
y in xis substring of"b" in "abc" is true
a x nrepeat"ab" x 3 is "ababab"
a[i]character from start"abc"[1] is "b"
a[-i]character from end"abc"[-1] is "c"
a[:]no change"abc"[:] is "abc"
a[s:]cut off start"abc"[1:] is "bc"
a[:-s]cut off end"abc"[:-1] is "ab"
a[s]reverse"abc"[::-1] is "cba"

Character strings are converted to integers, byte and bit strings as necessary using the current encoding and escape rules. For example when using a sane encoding "z"-"a" is 25.

Other types can be converted to character strings by using the type str or by using the repr and format functions.

Character strings support indexing and slicing. This is explained in detail in section Slicing and indexing.

mystr   = "oeU"         ; character string constant
        .text 'it''s'   ; it's
        .word "ab"+1    ; conversion result is "bb" usually

        .text "text"[:2]     ; "te"
        .text "text"[2:]     ; "xt"
        .text "text"[:-1]    ; "tex"
        .text "reverse"[::-1]; "esrever"

Byte string constants

Byte strings are like character strings, but hold bytes instead of characters.

Quoted character strings prefixing by b, l, n, p, s, x or z characters can be used to create byte strings. The resulting byte string contains what .text, .shiftl, .null, .ptext and .shift would create. Direct hexadecimal entry can be done using the x prefix and z denotes a z85 encoded byte string. Spaces can be used between pairs of hexadecimal digits as a separator for better readability.

Byte string operators and functions
y .. xconcatenate stringsx"12" .. x"34" is x"1234"
y in xis substring ofx"34" in x"1234" is true
a x nrepeatx"ab" x 3 is x"ababab"
a[i]byte from startx"abcd12"[1] is x"cd"
a[-i]byte from endx"abcd"[-1] is x"cd"
a[:]no changex"abcd"[:] is x"abcd"
a[s:]cut off startx"abcdef"[1:] is x"cdef"
a[:-s]cut off endx"abcdef"[:-1] is x"abcd"
a[s]reversex"abcdef"[::-1] is x"efcdab"

Byte strings support indexing and slicing. This is explained in detail in section Slicing and indexing.

Other types can be converted to byte strings by using the type bytes.

        .enc "screen"   ;use screen encoding
mystr   = b"oeU"        ;convert text to bytes, like .text
        .enc "none"     ;normal encoding

        .text mystr     ;text as originally encoded
        .text s"p1"     ;convert to bytes like .shift
        .text l"p2"     ;convert to bytes like .shiftl
        .text n"p3"     ;convert to bytes like .null
        .text p"p4"     ;convert to bytes like .ptext

Binary data may be embedded in source code by using hexadecimal byte strings. This is more compact than using .byte followed by a lot of numbers. As expected 1 byte becomes 2 characters.

        .text x"fce2"   ;2 bytes: $fc and $e2 (big endian)

If readability is not a concern then the more compact z85 encoding may be used which encodes 4 bytes into 5 characters. Data lengths not a multiple of 4 are handled by omitting leading zeros in the last group.

        .text z"FiUj*2M$hf";8 bytes: 80 40 20 10 08 04 02 01

For data lengths of multiple of 4 bytes any z85 encoder will do. Otherwise the simplest way to encode a binary file into a z85 string is to create a source file which reads it using the line label = binary('filename'). Now if the labels are listed to a file then there will be a z85 encoded definition for this label.

Lists and tuples

Lists and tuples can hold a collection of values. Lists are defined from values separated by comma between square brackets [1, 2, 3], an empty list is []. Tuples are similar but are enclosed in parentheses instead. An empty tuple is (), a single element tuple is (4,) to differentiate from normal numeric expression parentheses. When nested they function similar to an array. Both types are immutable.

List and tuple operators and functions
y .. xconcatenate lists[1] .. [2] is [1, 2]
y in xis member of list2 in [1, 2, 3] is true
a x nrepeat[1, 2] x 2 is [1, 2, 1, 2]
a[i]element from start("1", 2)[1] is 2
a[-i]element from end("1", 2, 3)[-1] is 3
a[:]no change(1, 2, 3)[:] is (1, 2, 3)
a[s:]cut off start(1, 2, 3)[1:] is (2, 3)
a[:-s]cut off end(1, 2.0, 3)[:-1] is (1, 2.0)
a[s]reverse(1, 2, 3)[::-1] is (3, 2, 1)
*aconvert to argumentsformat("%d: %s", *mylist)
... op aleft fold... + (1, 2, 3) is ((1+2)+3)
a op ...right fold(1, 2, 3) - ... is (1-(2-3))

Arithmetic operations are applied on the all elements recursively, therefore [1, 2] + 1 is [2, 3], and abs([1, -1]) is [1, 1].

Arithmetic operations between lists are applied one by one on their elements, so [1, 2] + [3, 4] is [4, 6].

When lists form an array and columns/rows are missing the smaller array is stretched to fill in the gaps if possible, so [[1], [2]] * [3, 4] is [[3, 4], [6, 8]].

Lists and tuples support indexing and slicing. This is explained in detail in section Slicing and indexing.

mylist  = [1, 2, "whatever"]
mytuple = (cmd_e, cmd_g)

mylist  = ("e", cmd_e, "g", cmd_g, "i", cmd_i)
keys    .text mylist[::2]    ; keys ("e", "g", "i")
call_l  .byte <mylist[1::2]-1; routines (<cmd_e-1, <cmd_g-1, <cmd_i-1)
call_h  .byte >mylist[1::2]-1; routines (>cmd_e-1, >cmd_g-1, >cmd_i-1)

Although lists elements of variables can't be changed using indexing (at the moment) the same effect can be achieved by combining slicing and concatenation:

lst     := lst[:2] .. [4] .. lst[3:]; same as lst[2] := 4 would be

Folding is done on pair of elements either forward (left) or reverse (right). The list must contain at least one element. Here are some folding examples:

minimum = size([part1, part2, part3]) <? ...
maximum = size([part1, part2, part3]) >? ...
sum     = size([part1, part2, part3]) + ...
xorall  = list_of_numbers ^ ...
join    = list_of_strings .. ...
allbits = sprites.(left, middle, right).bits | ...
all     = [true, true, true, true] && ...
any     = [false, false, false, true] || ...

The range(start, end, step) built-in function can be used to create lists of integers in a range with a given step value. At least the end must be given, the start defaults to 0 and the step to 1. Sounds not very useful, so here are a few examples:

;Bitmask table, 8 bits from left to right
        .byte %10000000 >> range(8)
;Classic 256 byte single period sinus table with values of 0–255.
        .byte 128 + 127.5 * sin(range(256) * pi / 128)
;Screen row address tables
_       := $400 + range(0, 1000, 40)
scrlo   .byte <_
scrhi   .byte >_

Dictionaries

Dictionaries hold key and value pairs. Definition is done by collecting key:value pairs separated by comma between braces {"key":"value", :"default value"}.

Looking up a non-existing key is normally an error unless a default value is given. An empty dictionary is {}. This type is immutable. There are limitations what may be used as a key but the value can be anything.

Dictionary operators and functions
y .. xcombine dictionaries{1:2, 3:4} .. {2:3, 3:1} is {1:2, 2:3, 3:1}
x[i]value lookup{"1":2}["1"] is 2
x.isymbol lookup{.ONE:1, .TWO:2}.ONE is 1
y in xis a key1 in {1:2} is true
; Simple lookup
        .text {1:"one", 2:"two"}[2]; "two"
; 16 element "fader" table 1->15->12->11->0
        .byte {1:15, 15:12, 12:11, :0}[range(16)]
; Symbol accessible values. May be useful as a function return value too.
coords  = {.x: 24, .y: 50}
        ldx #coords.x
        ldy #coords.y

Code

Code holds the result of compilation in binary and other enclosed objects. In an arithmetic operation it's used as the numeric address of the memory where it starts. The compiled content remains static even if later parts of the source overwrite the same memory area.

Indexing and slicing of code to access the compiled content might be implemented differently in future releases. Use this feature at your own risk for now, you might need to update your code later.

Label operators and functions
a.bb member of alabel.locallabel
.b in aif a has symbol b.locallabel in label
a[i]element from startlabel[1]
a[-i]element from endlabel[-1]
a[:]copy as tuplelabel[:]
a[s:]cut off start, as tuplelabel[1:]
a[:-s]cut off end, as tuplelabel[:-1]
a[s]reverse, as tuplelabel[::-1]
mydata  .word 1, 4, 3
mycode  .block
local   lda #0
        .endblock

        ldx #size(mydata) ;6 bytes (3*2)
        ldx #len(mydata)  ;3 elements
        ldx #mycode[0]    ;lda instruction, $a9
        ldx #mydata[1]    ;2nd element, 4
        jmp mycode.local  ;address of local label

Addressing modes

Addressing modes are used for determining addressing modes of instructions.

For indexing there must be no white space between the comma and the register letter, otherwise the indexing operator is not recognized. On the other hand put a space between the comma and a single letter symbol in a list to avoid it being recognized as an operator.

Addressing mode operators
#immediate
#+signed immediate
#-signed immediate
( )indirect
[ ]long indirect
,bdata bank indexed
,ddirect page indexed
,kprogram bank indexed
,rdata stack pointer indexed
,sstack pointer indexed
,xx register indexed
,yy register indexed
,zz register indexed

Parentheses are used for indirection and square brackets for long indirection. These operations are only available after instructions and functions to not interfere with their normal use in expressions.

Several addressing mode operators can be combined together. Currently the complexity is limited to 4 operators. This is enough to describe all addressing modes of the supported CPUs.

Valid addressing mode operator combinations
#immediatelda #$12
#+signed immediatelda #+127
#-signed immediatelda #-128
#addr,#addrmovemvp #5,#6
addrdirect or relativelda $12 lda $1234 bne $1234
bit,addrdirect page bitrmb 5,$12
bit,addr,addrdirect page bit relative jumpbbs 5,$12,$1234
(addr)indirectlda ($12) jmp ($1234)
(addr),yindirect y indexedlda ($12),y
(addr),zindirect z indexedlda ($12),z
(addr,x)x indexed indirectlda ($12,x) jmp ($1234,x)
[addr]long indirectlda [$12] jmp [$1234]
[addr],ylong indirect y indexedlda [$12],y
#addr,bdata bank indexedlda #0,b
#addr,b,xdata bank x indexedlda #0,b,x
#addr,b,ydata bank y indexedlda #0,b,y
#addr,ddirect page indexedlda #0,d
#addr,d,xdirect page x indexedlda #0,d,x
#addr,d,ydirect page y indexedldx #0,d,y
(#addr,d)direct page indirectlda (#$12,d)
(#addr,d,x)direct page x indexed indirectlda (#$12,d,x)
(#addr,d),ydirect page indirect y indexedlda (#$12,d),y
(#addr,d),zdirect page indirect z indexedlda (#$12,d),z
[#addr,d]direct page long indirectlda [#$12,d]
[#addr,d],ydirect page long indirect y indexedlda [#$12,d],y
#addr,kprogram bank indexedjsr #0,k
(#addr,k,x)program bank x indexed indirectjmp (#$1234,k,x)
#addr,rdata stack indexedlda #1,r
(#addr,r),ydata stack indexed indirect y indexedlda (#$12,r),y
#addr,sstack indexedlda #1,s
(#addr,s),ystack indexed indirect y indexedlda (#$12,s),y
addr,xx indexedlda $12,x
addr,yy indexedlda $12,y

Direct page, data bank, program bank indexed and long addressing modes of instructions are intelligently chosen based on the instruction type, the address ranges set up by .dpage, .databank and the current program counter address. Therefore the ,d, ,b and ,k indexing is only used in very special cases.

The immediate direct page indexed #0,d addressing mode is usable for direct page access. The 8 bit constant is a direct offset from the start of actual direct page. Alternatively it may be written as 0,d.

The immediate data bank indexed #0,b addressing mode is usable for data bank access. The 16 bit constant is a direct offset from the start of actual data bank. Alternatively it may be written as 0,b.

The immediate program bank indexed #0,k addressing mode is usable for program bank jumps, branches and calls. The 16 bit constant is a direct offset from the start of actual program bank. Alternatively it may be written as 0,k.

The immediate stack indexed #0,s and data stack indexed #0,r accept 8 bit constants as an offset from the start of (data) stack. These are sometimes written without the immediate notation, but this makes it more clear what's going on. For the same reason the move instructions are written with an immediate addressing mode #0,#0 as well.

The immediate (#) addressing mode expects unsigned values of byte or word size. Therefore it only accepts constants of 1 byte or in range 0–255 or 2 bytes or in range 0–65535.

The signed immediate (#+ and #-) addressing mode is to allow signed numbers to be used as immediate constants. It accepts a single byte or an integer in range −128–127, or two bytes or an integer of −32768–32767.

The use of signed immediate (like #-3) is seamless, but it needs to be explicitly written out for variables or expressions (#+variable). In case the unsigned variant is needed but the expression starts with a negation then it needs to be put into parentheses (#(-variable)) or else it'll change the address mode to signed.

Normally addressing mode operators are used in expressions right after instructions. They can also be used for defining stack variable symbols when using a 65816, or to force a specific addressing mode.

param   = #1,s            ;define a stack variable
const   = #1              ;immediate constant
        lda #0,b          ;always "absolute" lda $0000
        lda param         ;results in lda #$01,s
        lda param+1       ;results in lda #$02,s
        lda (param),y     ;results in lda (#$01,s),y
        ldx const         ;results in ldx #$01
        lda #-2           ;negative constant, $fe

Uninitialized memory

There's a special value for uninitialized memory, it's represented by a question mark. Whenever it's used to generate data it creates a hole where the previous content of memory is visible.

Uninitialized memory holes without previous content are not saved unless it's really necessary for the output format, in that case it's replaced with zeros.

It's not just data generation statements (e.g. .byte) that can create uninitialized memory, but .fill, .align or address manipulation as well.

*       = $200          ;bytes as necessary
        .word ?         ;2 bytes
        .fill 10        ;10 bytes
        .align 64       ;bytes as necessary

Booleans

There are two predefined boolean constant variables, true and false.

Booleans are created by comparison operators (<, <=, !=, ==, >=, >), logical operators (&&, ||, ^^, !), the membership operator (in) and the all and any functions.

Normally in numeric expressions true is 1 and false is 0, unless the -Wstrict-bool command line option was used.

Other types can be converted to boolean by using the type bool.

Boolean values of various types
bitsAt least one non-zero bit
boolWhen true
bytesAt least one non-zero byte
codeAddress is non-zero
floatNot 0.0
intNot zero
strAt least one non-zero byte after translation

Types

The various types mentioned earlier have predefined names. These can used for conversions or type checks.

Built-in type names
addressAddress type
bitsBit string type
boolBoolean type
bytesByte string type
codeCode type
dictDictionary type
floatFloating point type
gapUninitialized memory type
intInteger type
listList type
strCharacter string type
tupleTuple type
typeType type

Bit and byte string conversions can take a second parameter to specify and exact size. Values which can fit in shorter space will be padded but longer ones give an error.

bits(<expression>[, <bit count>])
Convert to the specific number of bits. If the number of bits is negative then it's a signed.
bytes(<expression>[, <byte count>])
Convert to the specific number of bytes. If the number of bits is negative then it's a signed.
        .cerror type(var) != str, "Not a string!"
        .text str(year)   ; convert to string

Symbols

Symbols are used to reference objects. Regularly named, anonymous and local symbols are supported. These can be constant or re-definable.

Scopes are where symbols are stored and looked up. The global scope is always defined and it can contain any number of nested scopes.

Symbols must be uniquely named in a scope, therefore in big programs it's hard to come up with useful and easy to type names. That's why local and anonymous symbols exists. And grouping certain related symbols into a scope makes sense sometimes too.

Scopes are usually created by .proc and .block directives, but there are a few other ways. Symbols in a scope can be accessed by using the dot operator, which is applied between the name of the scope and the symbol (e.g. myconsts.math.pi).

Regular symbols

Regular symbol names are starting with a letter and containing letters, numbers and underscores. Unicode letters are allowed if the -a command line option was used. There's no restriction on the length of symbol names.

Care must be taken to not use duplicate names in the same scope when the symbol is used as a constant as there can be only one definition for them.

Duplicate names in parent scopes are not a problem and this gives the ability to override names defined in lower scopes. However this can just as well lead to mistakes if a lower scoped symbol with the same name was meant so there's a -Wshadow command line option to warn if such ambiguity exists.

Case sensitivity can be enabled with the -C command line option, otherwise all symbols are matched case insensitive.

For case insensitive matching it's possible to check for consistent symbol name use with the -Wcase-symbol command line option.

A regular symbol is looked up first in the current scope, then in lower scopes until the global scope is reached.

f       .block
g        .block
n        nop            ;jump here
         .endblock
        .endblock

        jsr f.g.n       ;reference from a scope
f.x     = 3             ;create x in scope f with value 3

Local symbols

Local symbols have their own scope between two regularly named code symbols and are assigned to the code symbol above them.

Therefore they're easy to reuse without explicit scope declaration directives.

Not all regularly named symbols can be scope boundaries just plain code symbol ones without anything or an opcode after them (no macros!). Symbols defined as procedures, blocks, macros, functions, structures and unions are ignored. Also symbols defined by .var, := or = don't apply, and there are a few more exceptions, so stick to using plain code labels.

The name must start with an underscore (_), otherwise the same character restrictions apply as for regular symbols. There's no restriction on the length of the name.

Care must be taken to not use the duplicate names in the same scope when the symbol is used as a constant.

A local symbol is only looked up in it's own scope and nowhere else.

incr    inc ac
        bne _skip
        inc ac+1
_skip   rts

decr    lda ac
        bne _skip
        dec ac+1
_skip   dec ac          ;symbol reused here
        jmp incr._skip  ;this works too, but is not advised

Anonymous symbols

Anonymous symbols don't have a unique name and are always called as a single plus or minus sign. They are also called as forward (+) and backward (-) references.

When referencing them - means the first backward, -- means the second backwards and so on. It's the same for forward, but with +. In expressions it may be necessary to put them into brackets.

        ldy #4
-       ldx #0
-       txa
        cmp #3
        bcc +
        adc #44
+       sta $400,x
        inx
        bne -
        dey
        bne --

Excessive nesting or long distance references create poorly readable code. It's also very easy to copy-paste a few lines of code with these references into a code fragment already containing similar references. The result is usually a long debugging session to find out what went wrong.

These references are also useful in segments, but this can create a nice trap when segments are copied into the code with their internal references.

        bne +
        #somemakro      ;let's hope that this segment does
+       nop             ;not contain forward references...

Anonymous symbols are looked up first in the current scope, then in lower scopes until the global scope is reached.

Anonymous labels within conditionally assembled code are counted even if the code itself is not compiled and the label won't get defined. This ensures that anonymous labels are always at the same "distance" independent of the conditions in between.

Constant and re-definable symbols

Constant symbols can be created with the equal sign. These are not re-definable. Forward referencing of them is allowed as they retain the objects over compilation passes.

Symbols in front of code or certain assembler directives are created as constant symbols too. They are bound to the object following them.

Re-definable symbols can be created by the .var directive or := construct. These are also called as variables. They don't carry their content over from the previous pass therefore it's not possible to use them before their definition.

If the variable already exists in the current scope it'll get updated. If an existing variable needs to be updated in a parent scope then the ::= variable reassign operator is able to do that.

Variables can be conditionally defined using the :?= construct. If the variable was defined already then the original value is retained otherwise a new one is created with this value.

WIDTH   = 40            ;a constant
        lda #WIDTH      ;lda #$28
variabl .var 1          ;a variable
var2    := 1            ;another variable
variabl .var variabl + 1;update it verbosely
var2    += 1            ;compound assignment (add one)
var3    :?= 5           ;assign 5 if undefined

The star label

The * symbol denotes the current program counter value. When accessed it's value is the program counter at the beginning of the line. Assigning to it changes the program counter and the compiling offset.

Built-in functions

Built-in functions are pre-assigned to the symbols listed below. If you reuse these symbols in a scope for other purposes then they become inaccessible, or can perform a different function.

Built-in functions can be assigned to symbols (e.g. sinus = sin), and the new name can be used as the original function. They can even be passed as parameters to functions.

Mathematical functions

floor(<expression>)
Round down. E.g. floor(-4.8) is -5.0
round(<expression>)
Round to nearest away from zero. E.g. round(4.8) is 5.0
ceil(<expression>)
Round up. E.g. ceil(1.1) is 2.0
trunc(<expression>)
Round down towards zero. E.g. trunc(-1.9) is -1
frac(<expression>)
Fractional part. E.g. frac(1.1) is 0.1
sqrt(<expression>)
Square root. E.g. sqrt(16.0) is 4.0
cbrt(<expression>)
Cube root. E.g. cbrt(27.0) is 3.0
log10(<expression>)
Common logarithm. E.g. log10(100.0) is 2.0
log(<expression>)
Natural logarithm. E.g. log(1) is 0.0
exp(<expression>)
Exponential. E.g. exp(0) is 1.0
pow(<expression a>, <expression b>)
A raised to power of B. E.g. pow(2.0, 3.0) is 8.0
sin(<expression>)
Sine. E.g. sin(0.0) is 0.0
asin(<expression>)
Arc sine. E.g. asin(0.0) is 0.0
sinh(<expression>)
Hyperbolic sine. E.g. sinh(0.0) is 0.0
cos(<expression>)
Cosine. E.g. cos(0.0) is 1.0
acos(<expression>)
Arc cosine. E.g. acos(1.0) is 0.0
cosh(<expression>)
Hyperbolic cosine. E.g. cosh(0.0) is 1.0
tan(<expression>)
Tangent. E.g. tan(0.0) is 0.0
atan(<expression>)
Arc tangent. E.g. atan(0.0) is 0.0
tanh(<expression>)
Hyperbolic tangent. E.g. tanh(0.0) is 0.0
rad(<expression>)
Degrees to radian. E.g. rad(0.0) is 0.0
deg(<expression>)
Radian to degrees. E.g. deg(0.0) is 0.0
hypot(<expression y>, <expression x>)
Polar distance. E.g. hypot(4.0, 3.0) is 5.0
atan2(<expression y>, <expression x>)
Polar angle in −pi to +pi range. E.g. atan2(0.0, 3.0) is 0.0
abs(<expression>)
Absolute value. E.g. abs(-1) is 1
sign(<expression>)
Returns the sign of value as −1, 0 or 1 for negative, zero and positive. E.g. sign(-5) is -1

Byte string functions

These functions return byte strings of various lengths for signed numbers, unsigned numbers and addresses.

The naming of functions is not a coincidence and they return the bytes what the data directives with the same names normally emit.

byte(<expression>)
char(<expression>)
Return a single byte string from a 8 bit unsigned (0–255) or signed number (−128–127). E.g. byte(0) is x"00" and char(-1) is x"ff"
word(<expression>)
sint(<expression>)
Return a little endian byte string of 2 bytes from a 16 bit unsigned (0–65535) or signed number (−32768–32767). E.g. word(1024) is x"0004" and sint(-1) is x"ffff"
long(<expression>)
lint(<expression>)
Return a little endian byte string of 3 bytes from a 24 bit unsigned (0–16777216) or signed number (−8388608–8388607). E.g. long(123456) is x"40E201" and lint(-1) is x"ffffff"
dword(<expression>)
dint(<expression>)
Return a little endian byte string of 4 bytes from a 32 bit unsigned (0–4294967296) or signed number (−2147483648–2147483647). E.g. dword(123456789) is x"15CD5B07" and dint(-1) is x"ffffffff"
addr(<expression>)
Return a little endian byte string of 2 bytes from an address in the current program bank. E.g. addr(start) is x"0d08"
rta(<expression>)
Return a little endian byte string of 2 bytes from a return address in the current program bank. E.g. rta(4096) is x"ff0f"

Other functions

all(<expression>)
Return truth for various definitions of all.
All function
all bits set or no bits at allall($f) is true
all characters non-zero or empty stringall("c") is true
all bytes non-zero or no bytesall(x"ac24") is true
all elements true or empty listall([true, true, false]) is false

Only booleans in a list are accepted with the -Wstrict-bool command line option.

any(<expression>)
Return truth for various definitions of any.
Any function
at least one bit setany(~$f) is false
at least one non-zero characterany("c") is true
at least one non-zero byteany(x"ac24") is true
at least one true elementany([true, true, false]) is true

Only booleans in a list are accepted with the -Wstrict-bool command line option.

binary(<string expression>[, <offset>[, <length>]])
Returns the binary file content as bytes.

This function reads the content of a binary file as a byte string. It also accepts optional offset and length parameters.

Binary function invocation types
Read everythingbinary(name)
Skip starting bytesbinary(name, offset)
Some bytes from offsetbinary(name, offset, length)
sid     = binary("music.sid"); read in the SID file as bytes
offs    := sid[[$7, $6]]     ; data offset (big endian)
load    := sid[[$9, $8]]     ; load address (big endian)
init    = sid[[$b, $a]]      ; init address (big endian)
play    = sid[[$d, $c]]      ; play address (big endian)

; if load address is zero then it's the first 2 bytes of data
        .if load == 0
load    := sid[offs:offs+2]  ; load address (little endian)
offs    += 2                 ; skip load address bytes
        .endif

*       = load               ; set pc to load address
        .text sid[offs:]     ; dump music data
format(<string expression>[, <expression>, …])
Create string from values according to a format string.

The format function converts a list of values into a character string. The converted values are inserted in place of the % sign. Optional conversion flags and minimum field length may follow, before the conversion type character. These flags can be used:

Formatting flags
#alternate form (-$a, ~$a, -%10, ~%10, -10.)
*width/precision from list
.precision
0pad with zeros
-left adjusted (default right)
 blank when positive or minus sign
+sign even if positive
~binary and hexadecimal as bits

The following conversion types are implemented:

Formatting conversion types
bbinary
cUnicode character
ddecimal
e Eexponential float (uppercase)
f Ffloating point (uppercase)
g Gexponential/floating point
sstring
rrepresentation
x Xhexadecimal (uppercase)
%percent sign
        .text format("%#04x bytes left", 1000); $03e8 bytes left
len(<expression>)
Returns the number of elements.
Length of various types
bit stringlength in bitslen($034) is 12
character stringnumber of characterslen("abc") is 3
byte stringnumber of byteslen(x"abcd23") is 3
tuple, listnumber of elementslen([1, 2, 3]) is 3
dictionarynumber of elementslen({1:2, 3:4]) is 2
codenumber of elementslen(label)
random([<expression>, …])
Returns a pseudo random number.

The sequence does not change across compilations and is the same every time. Different sequences can be generated by seeding with .seed.

Random function invocation types
floating point number 0.0 <= x < 1.0random()
integer in range of 0 <= x < erandom(e)
integer in range of s <= x < erandom(s, a)
integer in range of s <= x < e, step trandom(s, a, t)
        .seed 1234      ; default is boring, seed the generator
        .byte random(256); a pseudo random byte (0–255)
        .byte random([16] x 8); 8 pseudo random bytes (0–15)
range(<expression>[, <expression>, …])
Returns a list of integers in a range, with optional stepping.
Range function invocation types
integers from 0 to e-1range(e)
integers from s to e-1range(s, a)
integers from s to e (not including e), step trange(s, a, t)
        .byte range(16) ; 0, 1, ..., 14, 15
        .char range(-5, 6); -5, -4, ..., 4, 5
mylist  = range(10, 0, -2); [10, 8, 6, 4, 2]
repr(<expression>)
Returns a string representation of value.
        .warn repr(var) ; pretty print value, for debugging
size(<expression>)
Returns the size of code, structure or union in bytes.
var     .word 0, 0, 0
        ldx #size(var)  ; 6 bytes
var2    = var + 2       ; start 2 bytes later
        ldx #size(var2) ; what remains is 4 bytes
sort(<list expression>)
Returns a sorted list or tuple.

If the original list contains further lists then these must be all of the same length. In this case the order of lists is determined by comparing their elements from the start until a difference is found. The sort is stable.

; sort IRQ routines by their raster lines
sorted  = sort([(60, irq1), (50, irq2)])
lines   .byte sorted[:, 0] ; 50, 60
irqs    .addr sorted[:, 1] ; irq2, irq1

Expressions

Operators

The following operators are available. Not all are defined for all types of arguments and their meaning might slightly vary depending on the type.

Unary operators
-negative +positive
!not ~invert
*convert to arguments ^decimal string

The ^ decimal string operator will be changed to mean the bank byte soon. Please update your sources to use format("%d", xxx) instead! This is done to be in line with it's use in most other assemblers.

Binary operators
+add -subtract
*multiply /divide
%modulo **raise to power
|binary or ^binary xor
&binary and <<shift left
>>shift right .member
..concat xrepeat
incontains !inexcludes

Spacing must be used for the x and in operators or else they won't be recognized as such. For example the expression [1,2]x2 should be written as [1,2]x 2 instead.

Parenthesis (( )) can be used to override operator precedence. Don't forget that they also denote indirect addressing mode for certain opcodes.

        lda #(4+2)*3

Comparison operators

Traditional comparison operators give false or true depending on the result.

The compare operator (<=>) gives −1 for less, 0 for equal and 1 for more.

Comparison operators
<=>compare   
==equals !=not equal
<less than >=more than or equals
>more than <=less than or equals
===identical !==not identical

Bit string extraction operators

These unary operators extract 8 or 16 bits. Usually they are used to get parts of a memory address.

Bit string extraction operators
<lower byte >higher byte
<>lower word >`higher word
><lower byte swapped word `bank byte
        lda #<label     ; low byte of address
        ldy #>label     ; high byte of address
        jsr $ab1e

        ldx #<>source   ; word extraction
        ldy #<>dest
        lda #size(source)-1
        mvn #`source, #`dest; bank extraction

Please note that these prefix operators are not strongly binding like negation or inversion. Instead they apply to the whole expression to the right. This may be unexpected but is required for compatibility with old sources which expect this behaviour.

        lda #<label+10  ;This is <(label+10) and not (<label)+10

;The check below is wrong and should be written as (>start) != (>end)
        .cerror >start != >end;Effectively this is >(start != (>end))

Conditional operators

Boolean conditional operators give false or true or one of the operands as the result.

Logical and conditional operators
x || yif x is true then x otherwise y
x ^^ yif both false or true then false otherwise x || y
x && yif x is true then y otherwise x
!xif x is true then false otherwise true
c ? x : yif c is true then x otherwise y
c ?? x : yif c is true then x otherwise y (broadcasting)
x <? yif x is smaller then x otherwise y
x >? yif x is greater then x otherwise y
;Silly example for 1=>"simple", 2=>"advanced", else "normal"
        .text MODE == 1 && "simple" || MODE == 2 && "advanced" || "normal"
        .text MODE == 1 ? "simple" : MODE == 2 ? "advanced" : "normal"
;Limit result to 0 .. 8
light   .byte 0 >? range(-16, 101)/6 <? 8

Please note that these are not short circuiting operations and both sides are calculated even if thrown away later.

With the -Wstrict-bool command line option booleans are required as arguments and only the ? operator may return something else.

Address length forcing

Special addressing length forcing operators in front of an expression can be used to make sure the expected addressing mode is used. Only applicable when used directly at the mnemonic.

Address size forcing
@bto force 8 bit address
@wto force 16 bit address
@lto force 24 bit address (65816)
        lda @w $0000    ; force the use of 2 byte absolute addressing
        bne @b label    ; prevent upgrade to beq+jmp with long branches in use
        lda @w #$00     ; use 2 bytes independent of accumulator size

Compound assignment

These assignment operators are short hands for updating variables. Constants can't be changed of course.

The variables on the left must be defined beforehand by := or .var.

Compound assignment operators can modify variables defined in parent scopes as well.

Compound assignments
+=add -=subtract
*=multiply /=divide
%=modulo **=raise to power
|=binary or ^=binary xor
&=binary and ||=logical or
&&=logical and <<=shift left
>>=shift right ..=concat
<?=smaller >?=greater
x=repeat .=member
v       += 1            ; same as 'v ::= v + 1'

Slicing and indexing

Lists, character strings, byte strings and bit strings support various slicing and indexing possibilities through the [] operator.

Indexing elements with positive integers is zero based. Negative indexes are transformed to positive by adding the number of elements to them, therefore −1 is the last element. Indexing with list of integers is possible as well so [1, 2, 3][(-1, 0, 1)] is [3, 1, 2].

Slicing is an operation when parts of sequence is extracted from a start position to an end position with a step value. These parameters are separated with colons enclosed in square brackets and are all optional. Their default values are [start:maximum:step=1]. Negative start and end characters are converted to positive internally by adding the length of string to them. Negative step operates in reverse direction, non-single steps will jump over elements.

This is quite powerful and therefore a few examples will be given here:

Positive indexing a[x]
It'll simply extracts a numbered element. It is zero based, therefore "abcd"[1] results in "b".
Negative indexing a[-x]
This extracts an element counted from the end, −1 is the last one. So "abcd"[-2] results in "c".
Cut off end a[:to]
Extracts a continuous range stopping before to. So [10,20,30,40][:-1] results in [10,20,30].
Cut off start a[from:]
Extracts a continuous range starting from from. So [10,20,30,40][-2:] results in [30,40].
Slicing a[from:to]
Extracts a continuous range starting from element from and stopping before to. The two end positions can be positive or negative indexes. So [10,20,30,40][1:-1] results in [20,30].
Everything a[:]
Giving no start or end will cover everything and therefore results in a complete copy.
Reverse a[::-1]
This gives everything in reverse, so "abcd"[::-1] is "dcba".
Stepping through a[from:to:step]
Extracts every stepth element starting from from and stopping before to. So "abcdef"[1:4:2] results in "bd". The from and to can be omitted in case it starts from the beginning or end at the end. If the step is negative then it's done in reverse.
Extract multiple elements a[list]
Extract elements based on a list. So "abcd"[[1,3]] will be "bd".

The fun start with nested lists and tuples, as these can be used to create a matrix. The examples will be given for a two dimensional matrix for easier understanding, but this also works in higher dimensions.

Extract row a[x]
Given a [(1,2),(3,4)] matrix [0] will give the first row which is (1,2)
Extract row range a[from:to]
Given a [(1,2),(3,4),(5,6),(7,8)] matrix [1:3] will give [(3,4),(5,6)]
Extract column a[x]
Given a [(1,2),(3,4)] matrix [:,0] will give the first column of all rows which is [1,3]
Extract column range a[:,from:to]
Given a [(1,2,3,4),(5,6,7,8)] matrix [:,1:3] will give [(2,3