PCC2 HACKING GUIDE
******************

  Table of Contents:
  - Basic Concepts
  - Initialisation
  - Data Structures
  - Format Strings
  - Internationalisation
  - Memory Management
  - Graphics Primitives
  - Source Code Structure
  - Naming Conventions
  - Debugging the Script Compiler
  - Character Sets
  - Modules
  - Terms



Basic Concepts
==============

  One game per program instance. Therefore, there's one global instance
  of all spec data (config, etc.).

  There can be multiple turns loaded. Code in game/ always takes the
  respective turn as a parameter. Code in client/ uses getDisplayedTurn().



Initialisation
==============

  PCC2 consists of several modules which must all be initialized
  consistently. We'll be building several executables, each of which
  uses a different set of modules. We don't want to initialize more
  than needed to avoid enlarging the executables, but we must make
  sure that all modules are initialized consistently.

  To solve this, each module defines a global variable mod_XXX which
  describes the module. The main program defines a variable
  yyy_modules, which lists all these modules. Upon a state transition,
  the main program calls performInitialisation.

  The following state chart contains all permitted transitions. Some
  transitions can be combined; in particular, OnEnterDirectory and
  OnEnterGame can be performed at once.


                   .-------------------------------._
                  /                                  '->
  (*) ----> OnStartup ---------------> OnEnterNoGame(A)---> OnShutdown
                |               _.-'         A        ->
                |           _.-'             |      -'
                V        <-'                 |    -'
       OnEnterDirectory(B) ---------> OnLeftDirectory(B)
                |        <-._                A
                |            '-._            |
                V                '-._        |
         OnEnterGame(C) ----------------> OnLeftGame(C)
                |                            A
                | .o( load game )            | .o( save game )
                V                            |
        OnEnteredGame(D) --------------> OnLeaveGame(D)
                             .
                              o( play game )

  (*) Initial state
  (A) spec_file_dir points to root (=specs directory)
  (B) spec_file_dir points to game, game_file_dir is set
  (C) like (B), player Ids are set, game not loaded
  (D) like (C), game is loaded and fully accessible

  The script scripts/checkinits.pl takes a (Linux) binary and attempts
  to verify that all initialisation data structures are set up
  correctly.



Data Structures
===============

  For details on file formats, read the File Format List available on
     http://phost.de/~stefan/filefmt.html
  It documents both regular VGAP file formats as well as PCC file formats,
  valid for PCC 1.x and PCC2.

  Most data structures are defined in game/struct.in. The script
  scripts/iogen.pl generates a .h/.cc file pair with unpacking tables for
  routines in io/io.h.

  Traditionally, the classic data structures have been used all through the
  core of PCC2. I'm now moving away from that. This allows us to use
  storage formats that differ from classic *.dis/*.dat files (e.g. network
  play using c2server or planets.nu), and provides a much cleaner interface
  to history information.

  Guidelines for future code (and transition):
  - do not use NUM_SHIPS, NUM_PLANETS for iteration or range checking. Use
    univ.ty_any_planets, univ.ty_any_ships, etc. instead.
  - do not make functions that return references to raw structures. For
    writing files, make a function that fills in a structure instead.


Maybe types and Data Validity
-----------------------------

  Most properties of units are exported as Maybe types (mp16_t etc.). The
  names of these types are "mpXX_t" if the valid values always are positive
  (this is the rule; -1 is used as "unknown" marker), and "mnXX_t" if valid
  values can also be negative (this is the exception, used for waypoint
  offsets, happiness, and cargo amounts that can be negative due to
  overloading).

  Before using these, the value must be checked (isValid()).

  Objects have a playability. A playability other than NotPlayable
  indicates that all "classic" parameters of the unit are known. Thus, if a
  unit is playable, isValid() checks can be omitted.


Cargo Transactions
------------------

  Cargo transfers and conversions (=building stuff) are implemented as
  transactions that can possibly be long-running. If the transaction
  consumes cargo, that will be marked as reserved and unavailable to other
  transactions.

  The idea is to have the user set up a transaction and then commit it
  (e.g. cargo transfer or ship building dialog) while still allowing a
  script invoked inbetween to also manipulate cargo. In theory, we could
  allow two concurrent "Cargo Transfer" windows.

  An example where this is actually used is when you call the "sell
  supplies" function from the planetary structure build screen.



Format Strings
==============

  See comment at head of cpluslib/format.cc.

  Right now, the format routine can distinguish between integers being
  one or zero in order to generate plural forms. For example,
     This ship has %d engine%!1{s%}.
  turns into "This ship has 1 engine." or "This ship has 8 engines.". If
  your language needs to differentiate more cases (I recall slavic
  languages differentiate between "1", "2..4" and "5 or more"?), it's no
  problem to add new modifiers. This way, only the format routine has to
  be taught of new languages, the actual program code can be left alone.

  Another important property is that the format routine is "forgiving"
  against type errors:
     format("%d", numToString(12999))
  will work and set status flags as if it were formatting a number.



Internationalisation
====================

  See comment at head of util/nls.cc for details.

  We use a homegrown gettext implementation. The source format for language
  databases is compatible to GNU gettext so we can "steal" GNU's xgettext
  and msgmerge. It should even be possible to replace the homegrown version
  by GNU gettext if someone insists. However, only the homegrown version
  gives us full control over the file format.



Memory Management
=================

  Essentially, we have two schemes of memory management:

  + Smart Pointers.
    "Ptr<T>" is a reference-counted smart pointer. It supports almost
    everything a dumb pointer can. In particular, it can be passed
    through a dumb pointer, in which case it finds its reference count
    using a global std::map. Classes derived from Object have their
    reference count embedded in the object.

  + Explicit Management.
    Essentially, whenever you create an object you have to decide how
    you get rid of it.


Smart Pointers
---------------

  Pros:
  + simple. you can build almost every data structure and it (usually)
    works. That's the reason I originally chose this for many parts.
  + factory functions are easy.

  Cons:
  + syntax problems. In a function call `foo(new X)', you don't see
    whether the parameter is a smart pointer or you have a leak. Plus,
    `Ptr<X>' looks uglier than `X*' (IMHO).
  + no cycles. Back-links must be dumb pointers or you have a leak
    again. I played a bit with "weak" pointers, which increases the
    mess instead of reducing it.
  + no stack-allocated objects.
  + generates *lots* of code.

  Currently, smart pointers are used where factories and wild sharing
  are required, and in some code not (yet?) converted to explicit
  memory management:
  + stream subsystem (needs factories, maybe sharing, objects come and go)
  + resmgr subsystem (needs factories, maybe sharing, objects come and go)
  + parts of graphics subsystem (animations: needs factories, sharing,
    objects come and go)
  + parts of interpreter subsystem (IntBytecodeObject needs sharing)


Explicit Memory Management
---------------------------

  Whenever an object is created, the creator has to think about how
  this object is going to be destroyed, and how long it's going to
  live. There are several choices:
  + stack allocation
  + putting it into a container who explicitly owns it (ptr_vector<T>
    or ObjectHolder).

  Pros:
  + typically generates small code
  + supports heap-allocated objects, arrays

  Cons:
  + programmer has to ensure destruction

  A useful convention applied throughout much of the code is "whenever
  I give you a reference parameter, I ensure the object lives as long
  as it needs." (which may be longer than during the function call).
  Since "new" returns a pointer, this ensures people think about what
  to do with their heap-allocated objects, because a mere "add(new
  Foo())" does not work.

  Unlike in the smart pointer case, you must make sure that child
  objects don't die before their parents when building object trees.
  For UIWidget, children nicely say good-bye before dying and remove
  themselves for their parent, so this is not a problem for widgets.
  For simple dialogs, it is simplest to explicitly allocate everything
  on the stack.

  Another useful tool to help with explicit MM is ObjectHolder.
  Instead of creating many named objects, you can also allocate
  objects on the heap and register them in an ObjectHolder. The
  ObjectHolder will destroy all its contents when it dies itself. In
  particular, "ObjectHolder.add" takes a pointer-to-T and returns a
  reference-to-T, so it suits as a nice adapter between `new' and
  parameter passing. Remember to keep the time between allocation and
  handing it to an ObjectHolder short, for exception safety.
  "ObjectHolder.add" is exception-safe.

  Factories are also possible: whenever you ask someone to give you an
  object, you must pass them an ObjectHolder, too. If the factory
  function needs to dynamically allocate new objects, it can use the
  ObjectHolder. This is used, for example, in UIBaseWidget::getCanvasFor().

  Explicit memory management is used in
  + game subsystem. GGameTurn is a container for quite a lot of stuff,
    and hands out references.
  + UI subsystem.
  + interpreter subsystem.

  Whenever possible, objects are passed by reference, not (smart)
  pointer.



Graphics Primitives
====================

  The most basic element is a GfxCanvas, which implements basic operations
  that could in theory be implemented in assembly. The concrete instance is
  a GfxSurface (an SDL_Surface, although I don't really like that so much
  SDL specific shines through -- I originally wanted this to be
  retargetable). Other GfxCanvas descendants implement filters.

  Although most routines that draw take a GfxCanvas, drawing usually uses a
  GfxContext, which is a GfxCanvas plus some high-level parameters ("line
  width").


Pitfalls
---------

  Some routines use half-open intervals, some use closed intervals. Be
  careful.

  Fonts use simple, brute-force anti-aliasing, that is, they draw a few
  pixels in half intensity. Therefore, make sure that text is never drawn
  twice without drawing background inbetween. (This rule is probably not
  followed everywhere, because traditionally only the big font, FONT_TITLE,
  has used anti-aliasing, so it didn't matter for our bread-and-butter
  font, FONT_NORMAL).



Source Code Structure
=====================

  Directory      | Type   | Content
  ================================================================
  P9             | source | input files for makefile generator
  ---------------+--------+---------------------------------------
  arch           | platf  | platform-specific code
    unix         |        | - Unix
    win32*       |        | - Win32, various compilers
  ---------------+--------+---------------------------------------
  client         | pcc1   | client source code
    actions      |        | - actions
    dialogs      |        | - dialogs
    screens      |        | - screens
    tiles        |        | - tiles
    widgets      |        | - widgets
  ---------------+--------+---------------------------------------
  cpluslib       | cplus  | miscellaneous
  ---------------+--------+---------------------------------------
  doc            | doc    | developer and user documentation
  ---------------+--------+---------------------------------------
  flak           | source | Fleet Action Kombat add-on source
  ---------------+--------+---------------------------------------
  game           | pcc2   | game library
    actions      |        | - complex actions
  ---------------+--------+---------------------------------------
  gfx            | pcc1   | graphics primitives
  ---------------+--------+---------------------------------------
  installer      | source | Windows installer source code (NSIS)
  ---------------+--------+---------------------------------------
  int            | pcc2   | script interpreter
    if           |        | - interpreter interface to game/client
  ---------------+--------+---------------------------------------
  io             | pcc2   | file I/O
  ---------------+--------+---------------------------------------
  phost          | pcc2   | config definitions from PHost/PDK
  ---------------+--------+---------------------------------------
  plugins        | -      | control files to build plugins
  ---------------+--------+---------------------------------------
  po             | source | translations
  ---------------+--------+---------------------------------------
  reshack        | source | resource editor
  ---------------+--------+---------------------------------------
  resmgr         | pcc1   | resource manager
  ---------------+--------+---------------------------------------
  resource       | data   | data files (images, fonts, ...)
  ---------------+--------+---------------------------------------
  scripts        | source | utility scripts (perl, shell)
  ---------------+--------+---------------------------------------
  sound          | pcc1   | sound playback
  ---------------+--------+---------------------------------------
  specs          | data   | data files (standard ship list)
  ---------------+--------+---------------------------------------
  sstr           | source | <sstream> replacement by Dietmar Khl,
                 |        | posted on de.comp.lang.c++, needed for
                 |        | g++ <3.0.
  ---------------+--------+---------------------------------------
  tests          | source | various test programs
  ---------------+--------+---------------------------------------
  tools          | source | command-line tool source code
  ---------------+--------+---------------------------------------
  u              | source | unit tests
  ---------------+--------+---------------------------------------
  ui             | pcc1   | user interface kernel
    widgets      |        | - widgets
  ---------------+--------+---------------------------------------
  util           | pcc2   | miscellaneous


  pcc1       Part of libpcc1.a / pcc1.lib, GUI code
  pcc2       Part of libpcc2.a / pcc2.lib, non-GUI code
  platf      Part of libplatf.a / platf.lib, platform-specific code
  cplus      Part of libcplus.a / cplus.lib
  source     Additional source code
  data, doc  Additional files

  The source code has been structured such that a console program only
  needs to reference pcc2, platf and cplus, but not pcc1.

  The Makefiles are generated from the files in P9/ using a Makefile
  generator (proj9) maintained separately. The files in cplus are also
  maintained separately. The files in sstr/ are used only when the
  platform does not provide its own stringstream class (this was an
  issue when PCC2 was started; it's probably not needed any longer
  today).

  Platform-specific code is in arch/. The configure script will set up
  the directory such that including "arch/platform.h" will refer to
  the right file.

  Some file names are pretty cryptic because I tried to roughly fit
  them into 8.3, to keep the way open for a MSDOS/DJGPP version. This
  is no longer a prime goal. However, file names must still be unique
  within a library, i.e. there cannot be a ui/init.o and a gfx/init.o
  because both would end up as libpcc1.a(init.o).

  Originally, there was no subdirectory structure beyond the first
  level. Instead, items were structured by file name prefixes
  (e.g. "wfoo.cc" for widgets, "dlg-foo.cc" for dialogs), and some
  files contained a whole boatload of items. They are now split much
  finer, and in subdirectories (i.e. "client/widgets/foo.cc"), but
  still much old code remains; I generally move it when a major
  change happens.


Makefiles
---------

  The Makefile is generated from the stuff in P9/ by a Perl script. I
  wrote that script when Makefile generators were not as abundand as
  they are now. Its main objective is creating hierarchical projects
  (libraries + executables) and makefiles for multiple toolchains with
  little reliance on makefile magic.

  This script is available from the same place as PCC2; changes to the
  makefiles should be done in the input files instead.



Naming Conventions
==================

  FooBar           types
    TFooBar          POD data type
    GFooBar          game/ class
    WFooBar          client/ class
    GfxFooBar        gfx/ class
    IntFooBar        int/ class
    UIFooBar         ui/ class
    FooException     exception

  foo_t            typedef

  FOO_BAR          constant, macro

  foo_Bar          constant (enum)

  foo_bar          local variables, members, STL-like types

  fooBar()         functions; sometimes also used for locals



Debugging the Script Compiler
=============================

  PCC2 compiles scripts into an internal "bytecode" representation before
  executing it (where "bytecode" isn't exactly true; it's actually a "32-bit
  word code").

  To debug this, we have the 'teststmt' program.
  + entered commands are compiled, bytecode listed, and executed.
  + '!file' compiles a file, lists bytecode, and executes it.
  + '.compile FILE' compiles a script into an object file.
  + '.load FILE' loads an object file and executes it.
  + '?SUB' disassembles a subroutine's bytecode.
  + '.opt SUB' optimizes a subroutine and list its bytecode.

  PCC2 beta10 cannot deal with object files, but I plan offering support
  sometime later. Note that the object file format is currently not set
  in stone.

  To generate object files with specific content, there is an assembler
  'scripts/c2asm.pl', which is also used together with 'teststmt' to test the
  optimizer ('tests/test_opt.{pl,qs}').

  To dump object files or SCRIPTx.CC files, use 'scripts/c2vmdump.pl'.

  A quick overview about all instructions can be found in 'int/opcode.h'.



Character Sets
===============

  Command-line character set, file system character set: those appear
  only within the platform implementation, and are immediately
  converted to UTF-8 before use.
    - Windows NT: file names are translated between Unicode and UTF-8.
    - Windows 9x: file names are translated between "ANSI" and UTF-8,
      usually with a Unicode step inbetween. For both kinds, the "ANSI"
      command line is ignored and re-parsed from the Unicode version.
    - Unix: file names and command line are translated between the
      locale character set and UTF-8. Although many instances of bad
      encoding (e.g. Latin-1 file name in an UTF-8 locale) will work,
      it's possible that PCC2 will be unable to access files with
      invalid (e.g. non-minimally encoded) names.

  Internal character set, font character set: always UTF-8.

  Game character set: in data structures, always external character
  set, so we have to convert on every use.

  For FLAK (PDK-based), the internal character set is the same as the
  command-line/file system character set, as the PDK expects. Because
  translations always are UTF-8, player messages are translated from
  UTF-8 into player-specific game character sets as needed, which are
  derived from the players' language configuration.


Unicode Usage
--------------

  E000..E07F      - error characters. Used by Unix filesystem operation
                    to encode bytes 80..FF which the libc cannot decode,
                    to allow them to survive UTF-8 coding.

  E100..E10F      - replacement glyphs, upper-left
  E110..E11F      - replacement glyphs, upper-right
  E120..E12F      - replacement glyphs, lower-left
  E130..E13F      - replacement glyphs, lower-right

  E140...         - specials

  Replacement glyphs are used to display characters for which a font
  doesn't have a glyph. For Unicode characters, four of them are
  displayed atop each other; one from each range for the four hex
  digits from the character code. For error characters, an upper-left
  and a lower-right glyph are used. Fonts either have all 64 of these
  characters, with the appropriate hex digits, or just the first
  group, all producing an identical replacement character.

  The fonts bundled with PCC2 support all characters from the
  supported codepages, yielding for Western European/American, Eastern
  European, and Cyrillic alphabets. The font renderer assumes
  precomposed characters (NFC, used on Windows), and cannot handle
  decomposed ones (NFD, used on MacOS). It also doesn't support right-
  to-left writing.

  The following table lists the correspondence between PCC 1.x's
  codepage 437 variant and the Unicode codes used in PCC2.

    Character            PCC1/dec  Unicode  UTF8/hex
    -------------------  --------  -------  --------
    up arrow               30, 24   U+2191  E2 86 91
    down arrow             31, 25   U+2193  E2 86 93
    right arrow                26   U+2191  E2 86 92
    left arrow                 27   U+2190  E2 86 90
    left/right arrow           29   U+2194  E2 86 94
    up/down arrow              18   U+2195  E2 86 95
    page up arrow              24   U+219F  E2 86 9F
    page down arrow            25   U+21A1  E2 86 A1
    top arc CW arrow           29   U+E140  EE 85 80
    left triangle         17, 221   U+25C0  E2 97 80
    right triangle        16, 222   U+25B6  E2 96 B6
    thin bar                   18   U+E144  EE 85 84
    round bullet                7   U+2022  E2 80 A2
    square bullet             254   U+25A0  E2 96 A0
    en dash                    22   U+2013  E2 80 93
    corresponds to            158   U+2259  E2 89 99
    times                     159   U+00D7  C3 97
    greater or equal          242   U+2265  E2 89 A5
    less or equal             243   U+2264  E2 89 A4
    registered                169   U+00AE  C2 AE
    trademark                 170   U+2122  E2 84 A2
    ornament left             156   U+E142  EE 85 82
    ornament right            157   U+E143  EE 85 83
    middle dot                250   U+00B7  C2 B7
    check mark (square root)  251   U+2713  E2 9C 93



Modules (Historic)
===================

  The packages "c2server" (PCC2 application server core) and
  "planetscentral" (planetscentral.com microservices) were intended to
  be built within the PCC2 source tree.

  These packages are no longer maintained and may no longer compile.
  The server code is now part of PCC2NG.


Terms
=====

  "atom"          - a number that represents a string, used to implement
                    keymaps and marker tags in PCC(2).

  "AVCs"          - additional visual contacts (targets above 50 limit)

  "BASIC string"  - character string occupying a fixed number of bytes,
                    padded with spaces. In BASIC, declared as "STRING*n",
                    with n being the number of bytes. PCC2 also accepts
                    null termination, see file format list.

  "battle"        - multiple fights (VCRs) to resolve a conflict.

  "BCO"           - bytecode object, basic data structure storing a
                    compiled script.

  "combat"        - single fight (VCR), part of a battle.

  "config"        - game configuration (hconfig, pconfig).

  "fusion"        - a bytecode optimisation that executes multiple
                    instructions in one step, avoiding intermediate values.

  "NVC"           - non-visual contact, i.e. ship listed in SHIPXY file, but
                    not in SHIP or TARGET file.

  "object"        - game object, usually something that can appear on the map.

  "object type"   - container of similar ->objects, i.e. "all ships".

  "Pascal string" - character string of up to 255 bytes, stored with
                    preceding length byte.

  "pref"          - user preferences, i.e. client configuration.

  "rune"          - an UTF-8 encoded character consisting of multiple bytes.

  "spec"          - specification data (hullspec, beamspec, etc.) including
                    additional specification data used by PHost such as
                    hullfunc.txt.

  "target"        - visual contact, i.e. ship listed in TARGET file.

  "tile"          - widget that can be displayed on a control screen.

  "widget"        - user interface control.



-eof-
