Tuesday, October 27, 2009

Project Goals











 < Day Day Up > 





Project Goals



After doing a great deal of work on various client projects, it seems like a good time to take a breather and work on a personal project. To that end, we'll work on enhancing the presentation of the sidebar links in a personal journal. Let's define some basic design directions and see where they take us.



  • We'll be starting with a page that already has some styles, so the menu's styles need to fit in with the presentation that already exists.

  • The links in the menu should be visually separated from one another; that is, we don't want a list of links with no separators or other visual effects.

  • We should come up with a design that makes the menu feel open and airy so that the links seem to be a part of the other content in the design.

  • We should come up with another design that encloses the links in a box or some other visual device that obviously separates them from the main content.



With these goals in mind, it's time to get set up and start styling!













     < Day Day Up > 



    12.4. ThreadWeaver



    12.4. ThreadWeaver


    ThreadWeaver is now one of the KDE 4 core libraries.
    It is discussed here because its genesis contrasts in many ways with that of the
    Akonadi project, and thus serves as an interesting comparison. ThreadWeaver
    schedules parallel operations. It was conceived at a time when it was
    technically pretty much impossible to implement it with the libraries used by
    KDE, namely Qt. The need for it was seen by a number of developers, but it took
    until the release of Qt 4 for it to mature and become mainstream. Today, it is
    used in major applications such as KOffice and KDevelop. It is typically applied
    in larger-scale, more complex software systems, where the need for concurrency
    and out-of-band processing becomes more pressing.


    ThreadWeaver is a job scheduler for
    concurrency. Its purpose is to manage and arbitrate resource usage in
    multi-threaded software systems. Its second goal is to provide application
    developers with a tool to implement parallelism that is similar in its approach
    to the way they develop their GUI applications. These goals are high-level, and
    there are secondary ones at a smaller scale: to avoid brute-force
    synchronization and offer means for cooperative serialization of access to data;
    to make use of the features of modern C++ libraries, such as thread-safe,
    implicit sharing, and signal-slot-connections across threads; to integrate with
    the application's graphical user interface by separating the processing elements
    from the delegates that represent them in the UI; to allow it to dynamically
    throttle the work queue at runtime to adapt to the current system load; to be
    simplistic; and many more.


    The ThreadWeaver library was
    developed to satisfy the needs of developers of event-driven GUI programs, but
    it turned out to be more generic. Because GUI programs are driven by a central
    event loop, they cannot process time-consuming operations in their main thread.
    Doing so would freeze the user interface until the operation is finished. In
    some windowing environments, the user interface can be drawn only from the main
    thread, or the windowing system itself is single-threaded. So a natural way of
    implementing responsive cross-platform GUI applications is to perform all
    processing in worker threads and update the user interface from the main thread
    when necessary. Surprisingly, the need for concurrency in user interfaces is
    rarely ever as obvious as it should be, although it has been emphasized for
    OS/2, Windows NT, and Solaris eons ago. Multithreaded programming is more
    complicated and requires a better understanding of how the written code actually
    functions. Multithreadings also seems to be a topic well understood by software
    architects and designers, and badly disseminated to software maintainers and
    less-experienced programmers. Also, some developers seem to think that most
    operations are fast enough to be executed synchronously, even reading from
    mounted filesystems, which is a couple of orders of magnitude slower than
    anything processed in the CPU. Such mistakes surface only under extraordinary
    circumstances, such as when the system is under heavy I/O load or, more
    commonly, a mounted filesystem has been put to sleep to save power
    or—heavens!—when, all of a sudden, the filesystem happens to be on the
    network.


    The following section will describe
    the architecture of the library along with its underlying concepts. At the end
    of the chapter, we will explore how it found its way into KDE 4.


    12.4.1. Introduction to
    ThreadWeaver: Or, How Complicated Can It Be to Load a File?


    To convince programmers to make use of
    concurrency, it needs to be conveniently available. Here is a typical example
    for an operation performed in a GUI program—loading a file into a memory buffer
    to process it and display the results. In an imperative program, where all
    individual operations are blocking, it is of little complexity:






    1. Check whether the file exists and is
      readable.




    2. Open the file for reading.




    3. Read the file contents into memory.




    4. Process them.




    5. Display the results.


    To be
    user-friendly, it is sufficient to print a progress message to the command line
    after every step (if the user requested verbose mode).


    In a GUI program, things appear in a
    different light because during all of these steps, it is necessary to be able to
    update the screen, and users expect a way to cancel the operation. Although it
    sounds unbelievable, even recent documentation for GUI toolkits mentions the
    "check for events occasionally" approach. The idea is to periodically check for
    events while processing the aforementioned steps in chunks and update or abort
    if necessary. A lot of care needs to be applied in this situation because the
    application state can change in unexpected ways. For example, the user might
    decide to close the program, unaware that the program checks for an event in the
    midst of a call stack of an operation. To put it shortly, this approach of
    polling has never worked very well and is generally out of fashion.


    A better approach is to use a thread
    (of course). But without a framework to help, this often leads to weird
    implementations as well. Since GUI programs are event-based, every step just
    listed starts with an event, and an event notifies its completion. In the C++
    world, signals are often used for the notification. Some programs look like
    this:






    1. The user requests to load the file,
      which triggers a handler method by a signal or event.




    2. The operation to open and load the file
      is started and connected to a second method for notification about its
      completion.




    3. In this method, processing the
      data is started, connected to a third method.




    4. The last method finally displays
      the results.


    This cascade of handler methods does
    not track the state of the operation very well and is usually error-prone. It
    also shows a lack of separation of operations and the view. Nevertheless, it is
    found in many GUI applications.


    This is exactly where ThreadWeaver is
    there to help. Using jobs, the implementation will look like this:






    1. The user requests to
      load the file, which triggers a handler method by a signal or
      event.




    2. In the handler method, the user
      creates a sequence of jobs (a sequence is a job container that executes its jobs
      in the order they were added). He adds a job to load the file and one to process
      its contents. The sequence object is a job itself and sends a signal when all
      its contained jobs are completed. Up to this point, no processing has taken
      place; the programmer only declared what needs to be done in what order. Once
      the whole sequence is set up, the user queues it into the application-global job
      queue (a lazy initialized singleton). The sequence is automatically executed by
      worker threads.




    3. When the done() signal of
      the sequence is received, the data is ready to be
      displayed.


    Two aspects are apparent. First, the
    individual steps are all declared in one go and then executed. This alone is a
    major relief to GUI programmers because it is a nonexpensive operation and can
    easily be performed in an event handler. Second, the usual issues of
    synchronization can largely be avoided by the simple convention that the queuing
    thread only touches the job data after it has been prepared. Since no worker
    thread will access the job data anymore, access to the data is serialized, but
    in a cooperative fashion. If the programmer wants to display progress to the
    user, the sequence emits signals after the processing of every individual job
    (signals in Qt can be sent across threads). The GUI remains responsive and is
    able to dequeue the jobs or request cancellation of processing.


    Since it is much easier to implement I/O
    operations this way, ThreadWeaver was quickly adopted by programmers. It solved
    a problem in a nice, convenient way.


    12.4.2. Core Concepts and
    Features


    In the previous example, job
    sequences have been mentioned. Let us look at what other constructs are provided
    in the library.


    Sequences are a specialized form of
    job collections. Job collections are containers that queue a set of jobs in an
    atomic operation and notify the program about the whole set. Job collections are
    composites, in the way that they are implemented as job classes themselves.
    There is only one queuing operation in ThreadWeaver: it takes a Job pointer.
    Composite jobs help keep the queue API minimal.


    Job sequences use dependencies to make sure the
    contained jobs are executed in the correct order. If a dependency is declared
    between two jobs, it means that the depending job can be executed only after its
    dependency has finished processing. Since dependencies can be declared in an m:n
    fashion, pretty much all imaginable control flows of depending operations (which
    are all directed graphs, since repetition of jobs is not allowed) can be modeled
    in the same declarative fashion. As long as the execution graph remains
    directed, jobs may even queue other jobs while being processed. A typical
    example is that of rendering a web page, where the anchored elements are
    discovered only once the text of the HTML document itself is processed. Jobs can
    then be added to retrieve and prepare all linked elements, and a final job that
    depends on all these preparatory jobs renders the page for display. Still, no
    mutex necessary. Dependencies are what distinguish a scheduling system like
    ThreadWeaver from mere tools for parallel processing. They relieve the
    programmer of thinking how to best distribute the individual suboperations to
    threads. Even with modern concepts such as futures, usually the programmer still
    needs to decide on the order of operations. With ThreadWeaver, the worker
    threads eagerly execute all possible jobs that have no unresolved dependencies.
    Since the execution of concurrent flow graphs is inherently undeterministic, it
    is very unlikely that a manually defined order is flexible enough to be the most
    efficient. A scheduler can adapt much better here. Computer scientists tend to
    disagree with this thesis, whereas economists, who are more used to analyzing
    stochastic systems, often support it.


    Priorities can be used to influence the order of execution as well.
    The priority system used is quite simple: of all integer priorities assigned to
    jobs in the queue, the highest ones are first handed to an available worker
    thread. Since the job base class is implemented in a way that allows writing
    decorators, changing a job's priority externally can be done by writing a
    decorator that bumps the priority without touching the job's implementation. The
    combination of priorities and dependencies can lead to interesting results, as
    will be shown later.


    Instead of relying on direct implementations of queueing behavior, ThreadWeaver
    uses queue policies. Queue policies do not immediately affect when and how a
    particular job is executed. Instead, they influence the order in which jobs are
    taken from the queue by the worker threads. Two standard implementations come
    with ThreadWeaver. One is the dependencies discussed earlier. The other is
    resource restrictions. Using resource restrictions, it can be declared that of a
    certain subset of all created jobs (for example, local filesystem I/O-expensive
    ones), only a certain amount can be executed at the same time. Without such a
    tool, it regularly happens that some subsystems get overloaded. Resource
    restrictions act much like semaphores in traditional threading, except that they
    do not block a calling thread, and instead simply mark a job as not yet
    executable. The thread that checked whether the job can be executed is then able
    to try to get another job to execute.


    Queue policies are assigned to jobs,
    and the same policy object can be assigned to many. As such, they are composed,
    and every job can be managed by any combination of the available policies.
    Inheriting specialized policy-driven job base classes would not have provided
    such flexibility. Also, this way, job objects that do not need any extra
    policies are in no way affected by a possible performance hit of evaluating the
    policies.


    12.4.3. Declarative Concurrency: A
    Thumbnail Viewer Example


    Another example explains how these different
    ThreadWeaver concepts play together. It uses jobs, job composites, resource
    restrictions, priorities, and dependencies, all to render wee little thumbnail
    images in a GUI program. Let us first look at what operations are required to
    implement this function, how they depend, and how the user expects to be
    presented with the results. The example is part of the ThreadWeaver source
    code.


    In this example, it is assumed that
    loading the thumbnail preview for a digital photo involves three operations: to
    load the raw file data from disk, to convert the raw data into an image
    representation without changing its size, and then to scale to the required size
    of the thumbnail. It is possible to argue that the second and third steps could
    be merged into one, but that is (a) not the point of the exercise (just like
    streamed loading of the image data) and (b) would impose the restriction that
    only image formats can be used where the drivers support scaling during load. It
    is also assumed that the files are present on the hard disk. Since the
    processing of each file does not influence or depend on the processing of any
    other, all files can be processed in parallel. The individual three steps to
    process one file need to be performed in sequence.


    But that is not all. Since this is an
    example with a graphical user interface, the expectations of the user have to be
    kept in mind. It is assumed that the user is interested in visual feedback,
    which also gives him the impression of progress. Image previews should be shown
    as soon as they are available, and progress information in the form of a
    reliable progress bar would be nice. The user also expects that the program will
    not grind his computer to a halt, for example by excessive I/O operations.


    Different ThreadWeaver tools can be
    applied to this problem. First of all, processing an individual file is a
    sequence of three job implementations. The jobs are quite generic and can be
    part of a toolbox of premade job classes available to the application. The jobs
    are (the class names match the ones in the example source code):




    • A FileLoaderJob, which loads a
      file on the file system into an in memory byte array



    • A QImageLoaderJob to convert
      the image's raw data into the typical representation of an in-memory image in Qt
      applications (providing the application access to all available image decoders
      in the framework or registered by the application)



    • A ComputeThumbNailJob, which
      simply scales the image to the wanted size of the preview


    All of those are added to a JobSequence,
    and each of these sequences is added to a JobCollection. The composite implementation of the
    collection classes allow for implementations that represent the original problem
    very closely and therefore feel somewhat natural and canonical to the
    programmer.


    This solves part one of the problem, the
    parallel processing of
    the different images. It could easily lead to other problems, though. With the
    given declaration of the problem to the ThreadWeaver queue, there is nothing
    that prevents it from loading all files at once and only then starting to
    process images. Although this is unlikely, we haven't told the system otherwise
    yet. To make sure that only so many file loaders are started at the same time, a
    resource restriction is used. The code for it looks like this:

    #include "ResourceRestrictionPolicy.h"

    ...

    static QueuePolicy* resourceRestriction()
    {
    static ResourceRestrictionPolicy policy( 4 );
    return &policy;
    }


    File loaders simply apply the policy
    in their constructor or, if generic classes are used, when they are created:

    fileloader->assignQueuePolicy( resourceRestriction() );


    But that still does not completely
    arrange the order of the execution of the jobs exactly as wanted. The queue
    might now start only four file loaders at once, but it still might load all the
    files and then calculate the previews (again, this is a very unlikely behavior).
    It needs one more tool, and a bit of thinking against the grain, to solve the
    problem, and this is where priorities come into play. The problem, translated
    into ThreadWeaver lingo, is that file loader jobs have lowest priority but need
    to be executed first; image loader jobs have precedence over file loaders, but a
    file loader must have finished first before an image loader can be started; and
    finally, thumbnail computer jobs have highest priority, even if they depend on
    the other two phases of processing. Since the three jobs are already in a
    sequence, which will make sure they are executed in the right order for every
    image, assigning priority one to file loaders, two to image loaders, and three
    to thumbnail computers finally solves the problem. Basically, the queue will now
    complete one thumbnail as soon as possible, but will not stop to load the images
    if slots for file loading become available. Since the problem is mostly I/O
    bound, this means that the total time until the thumbnails for all images are
    shown is a little more than the time it takes to load them from the hard disk
    (other factors aside, such as extremely high-resolution RAW images). In any
    sequential solution, the behavior would likely be much worse.


    The description of the solution might have
    felt complex, so lightening it up with a bit of code is probably in order. This
    is how the jobs are generated after the user has selected a couple of hundred
    images for processing:

    m_weaver->suspend();
    for (int index = 0; index < files.size(); ++index)
    {
    SMIVItem *item = new SMIVItem ( m_weaver, files.at(index ), this );
    connect ( item, SIGNAL( thumbReady(SMIVItem* ) ),
    SLOT ( slotThumbReady( SMIVItem* ) ) );
    }
    m_startTime.start();
    m_weaver->resume();


    To give correct progress feedback,
    processing is suspended before the jobs are added. Whenever a sequence is
    completed, the item object emits a signal to update the view. For every selected
    file, a specialized item is created, which in turn creates the job objects for
    processing one file:

    m_fileloader = new FileLoaderJob ( fi.absoluteFilePath(),  this );
    m_fileloader->assignQueuePolicy( resourceRestriction() );
    m_imageloader = new QImageLoaderJob ( m_fileloader, this );
    m_thumb = new ComputeThumbNailJob ( m_imageloader, this );
    m_sequence->addJob ( m_fileloader );
    m_sequence->addJob ( m_imageloader );
    m_sequence->addJob ( m_thumb );
    weaver->enqueue ( m_sequence );


    The priorities are virtual properties of
    the job objects and are set there. It is important to keep in mind that all
    these objects are set to not process until they are queued, and in this case,
    until the processing is explicitly resumed. So the whole operation to create all
    these sequences and jobs really takes only a very short time, and the program
    returns to the user immediately, for all practical matters. The view updates as
    soon as a preview image is available.


    12.4.4. From Concurrency to
    Scheduling: How to Implement Expected Behavior Systematically


    The previous
    examples have shown how analyzing the problem completely really helps to solve
    it (I hope this does not come as a surprise). To make sure concurrency is used
    to write better programs, it is not enough to provide a tool to move stuff to
    threads. The difference is scheduling: to be able to tell the program what
    operations have to be performed and in what order. The approach is remotely
    reminiscent of PROLOG programming lessons, and sometimes requires a similar way
    of thinking. Once the minds involved are sufficiently assimilated, the results
    can be very rewarding.


    One design decision of the
    central Weaver class has not been discussed yet.
    There are two very disjunct groups of users of the Weaver classes API. The internal Thread objects access it to retrieve their jobs to process, whereas
    programmers use it to manage their parallel operations. To make sure the public
    API is minimal, a combination of decorator and facade has been applied that
    limits the publicly exposed API to the functions that are intended to be used by
    application programmers. Further decoupling of the internal implementation and
    the API has been achieved by using the PIMPL idiom, which is generally applied
    to all KDE APIs.


    12.4.5. A Crazy Idea


    It has been mentioned earlier that
    at the time ThreadWeaver started to be developed, it was not really possible to
    implement all its ideas. One major obstacle was, in fact, a prohibitive one: the
    use of advanced implicit sharing features, which included reference counting, in
    the Qt library. Since this implicit sharing was not thread-safe, the passing of
    every simple Plain Old Data object (POD) was a synchronization point. The author
    assumed this to be impractical for users and therefore recommended against using
    the prototype developed with Qt 3 for any production environments. The
    developers of the KDEPIM suite (the same people who now develop Akonadi) thought
    they really knew better and immediately imported a preliminary ThreadWeaver
    version into KMail, where it is used to this day. Having run into many of the
    problems ThreadWeaver promised to solve, the KMail developers eagerly embraced
    it, willing to live with the shortcomings pointed out by its author, even
    against his express wishes.


    The fact that an imperfect version of the
    library was in active use in KDE served as a motivating factor for quickly porting it to
    Qt4 when that became usable in beta versions. Thus it was available rather early
    in the KDE 4 development cycle, if only in a secondary module and not yet as
    part of KDELibs. Over the course of two years, the author gave a number of
    presentations on the library, presenting an ever-easier and more complete API as
    he kept improving it. It was, one could say, a solution looking for a problem.
    The majority of developers working on KDE needed time to realize that this
    library was not only academic, but could improve their software significantly
    given that they make the investment in taking a step back and rethinking some of
    their architectural structures. There was no concrete need by a group of
    developers driving the library's progress; it was progressed by an individual
    because of his belief in the growing relevance of the problem and the importance
    of making available a good solution for the KDE 4 platform. Especially following
    the 2005 Akademy conference in Malaga, Spain, more programs started to use
    ThreadWeaver, including KOffice and KDevelop, which created enough momentum for
    it to be integrated into the main KDE 4 set of libraries.


    ThreadWeaver represents the
    case of an alternative solution to a problem that once it had matured to
    critical point and once the author and the prospective user community agreed
    that the time had come for it to be adopted by developers in their projects, it
    was quickly promoted to a cornerstone of KDE 4. After that, the attitudes of
    community members changed from mild amusement to appreciation and recognition of
    the effort that had gone into it. This is an example of how efficient this
    community can be at making technical decisions and adapting its stance when an
    approach proves itself in practice. There can be no doubt that ThreadWeaver is a
    much better library now than it would have been if it not taken three to four
    years of rubbing up against the KDE project until its inclusion. And this
    includes the rogue premature adoption by the KMail developers. There is also
    little doubt that applications written for KDE 4 can deal with concurrency a lot
    better and thus provide a better experience to their users, because it succeeded
    in the end.


    ThreadWeaver will be extended mostly by adding GUI components to visually represent queue activity, and by
    including more predefined job classes. Another idea is the integration with
    operating system IPC mechanisms (to allow for host-global resource restrictions,
    for example), but those are hindered by the requirement to be cross-platform.
    The approaches taken by the different operating systems are very diverse. With
    the public availability of the KDE 4 line, it became visible to a large
    audience. Since ThreadWeaver is not really KDE-specific, the question of where
    to go next (Freedesktop.org?) is in the air. For now, the focus remains to
    provide developers of applications and the desktop with a reliable scheduler for
    concurrency.


     


    Joining Strings










    Joining Strings






    print "Words:" + word1 + word2 + word3 + word4
    print "List: " + ' '.join(wordList)




    Strings can be joined together using a simple add operation, formatting the strings together or using the join() method. Using either the + or += operation is the simplest method to implement and start off with. The two strings are simply appended to each other.


    Formatting strings together is accomplished by defining a new string with string format codes, %s, and then adding additional strings as parameters to fill in each string format code. This can be extremely useful, especially when the strings need to be joined in a complex format.


    The fastest way to join a list of strings is to use the join(wordList) method to join all the strings in a list. Each string, starting with the first, is added to the existing string in order. The join method can be a little tricky at first because it essentially performs a string+=list[x] operation on each iteration through the list of strings. This results in the string being appended as a prefix to each item in the list. This actually becomes extremely useful if you want to add spaces between the words in the list because you simply define a string as a single space and then implement the join method from that string:


    word1 = "A"
    word2 = "few"
    word3 = "good"
    word4 = "words"
    wordList = ["A", "few", "more", "good", "words"]

    #simple Join
    print "Words:" + word1 + word2 + word3 + word4
    print "List: " + ' '.join(wordList)

    #Formatted String
    sentence = ("First: %s %s %s %s." %
    (word1,word2,word3,word4))
    print sentence

    #Joining a list of words
    sentence = "Second:"
    for word in wordList:
    sentence += " " + word
    sentence += "."
    print sentence


    join_str.py


    Words:Afewgoodwords
    List: A few more good words
    First: A few good words.
    Second: A few more good words.


    Output from join_str.py code












    Deriving Classes from Use Cases






























    Chapter 3 -
    Diagramming Business Objects
    byAndrew Filevet al.?
    Wrox Press ©2002

























    Team FLY






    Deriving Classes from Use Cases


    Now we're ready for the hard part - deriving business classes from use cases. Part of the problem is that use cases are not object-oriented. They simply represent a list of everything users can do with the computer system. We need to examine the description of the use case and derive business classes from it.


    When using techniques such as CRC cards (which are not part of the UML), we are encouraged to look at the nouns in the use case description as possible candidates for use cases. Here's a list of some of these key nouns (ignoring nouns that are obviously attributes of other entities such as Borrower ID and Media ID):




    • Librarian




    • Borrower




    • Media




    • Fines




    This process can actually get you pretty far, but I've found from experience it doesn't take you far enough. If you start out with this list of entities, you start going down a particular path, and have to backtrack and rework your object model. Although reworking or refactoring your model is part of the process, we can get ourselves closer to a working object model by thinking about data.





    Thinking about Data


    Why does thinking about our application's data help us in our object modeling efforts? As we mentioned earlier in this chapter, when we data model we often create tables representing real-world entities. Due to this relationship, you will often create a business object for each main table in your application, so thinking about data early is a wise decision.


    Although you want to start thinking about data, you don't want to get 'married' to a particular data model this early on. Use your data model to help you think things through, but don't set it in stone. Be willing to let your business object's behavior influence the data model. I have found data modeling helps shake the bugs out of an object model. Even if you wait to model data until after you've first tried to create your object model, you can test your model by creating test data, which you access from your business classes. In other cases, you may have a data structure you are forced to work with. In this case, you simply can't ignore the data structure. However, a good object model can help hide the flaws in an otherwise imperfect data model.


    So, let's start thinking about the structure of the data we need for our application. The following diagram shows a Visio data diagram containing tables we've started to flesh out for our library application:





    The first table in this diagram is the Borrower table. This is an easy place to start because our use case specifically mentions a Borrower ID. Notice the Borrower table has both a primary key field (BorrowerPK) and an ID field (ID). Most database designers agree it's important for all tables in your application to have a system-generated primary key used to uniquely identify each record in addition to any 'business' keys you may have. In this table, the ID field corresponds to the Borrower ID mentioned in the use case. This is the value manually entered by the Librarian or scanned from the Borrower's library card. However, the primary key is used when linking Borrower records to records in other tables. For good measure, we've also added a few obvious fields such as FirstName, LastName, and Email. We'll talk about the TrxLogFK foreign key field in just a bit.




    Media is another easy table. Obviously, we need to have a record of each media item borrowers can check out. Again, we have a situation where there is both a primary key field and an ID field. Again, we've added some obvious fields such as Title, Author, and Subject. There is also a foreign key pointer field (MediaTypeFK) to the MediaType table. Rather than storing the media type information directly in the Media table, we normalize our data by storing the information in a separate MediaType table.


    The MediaType table contains a description of the type of media (magazine, book, DVD), the number of days it can be checked out, and the daily fine if the item is overdue.


    The Transaction Log table (TrxLog) isn't as obvious. Our use cases specify we need to keep track of checked out media, whether or not it's overdue, as well as any unpaid fines. The easiest way to do this is to create a transaction log containing a record of each media check out/check in, including the media ID and the borrower ID. Although this information could be stored in the Media table, placing it there does not allow us to maintain a history, because these fields are overwritten the next time the media is checked out/in. The DueDate field provides a place where the media due date is stored. Although this information can be calculated dynamically, persisting it to the transaction record allows our system to account for any changes we make to the check out period rules. If the due date is calculated and saved when an item is checked out, even if the business rules change before the item is returned, we can still determine the correct due date for each piece of media.


    The Fine table contains fines applied to specific transaction log records. The TrxLogFK field is a foreign key pointer to the TrxLog table. The FineAmt field contains the amount of the fine and the PaymentAmt field contains any payment amount applied to this fine. I guarantee this accounting solution will bring tears to you financial wizards, but this simple solution works fine for our example.
















    Team FLY



    Chapter 7. Class Actions




    I l@ve RuBoard







    Chapter 7. Class Actions



    The declaration of a class on a class diagram doesn't "do" anything; the declaration merely states that when we create instances of the class, each object must have the data and behavior declared by the class.


    Actions do stuff: They create and delete objects, access attributes and links, make conditional choices, iterate, transform data, and otherwise generally compute. In the course of this book, we shall describe the actions you can specify that make the domain actually do something. In this chapter, we'll describe those actions that affect objects, links, and classes.



    Executable UML relies on the Precise Action Semantics for UML [1] adopted as an integral part of UML in late 2001. These action semantics provide for the specification of actions, but they do not define an action language syntax.


    Presently, therefore, there is no standard syntax for actions, though to specify the actions in an executable model, we have to use something, some concrete syntax. The syntax we use here is a real one, and it executes today [2]. The complete case study models for the online bookstore have been executed using this language, and the case study models are presented in Appendix B.



    Why NotJava?


    Why not just write Java? Or just use your favorite programming language?


    The answer has to do with raising the level of abstraction. To gain access to higher-level abstractions such as data structures and control structures, we gave up the ability to manipulate registers and the stack directly. This conferred independence from the hardware platform, in turn enabling portability of programs from one hardware platform to another.


    So it is with Executable UML and the action language. In return for the ability to work in terms of the domain objects directly, you give up pointer manipulation, arrays, lists, and various implementation tricks your language allows. This grants independence from the software platform. Now you can build Executable UML models and have them execute on a distributed system using CORBA, a small footprint embedded chip using C and no operating system, or a complex multi-processor implementation using C++�all without having to change the application models.


    To garner these benefits, the action semantics:



    • defines statements that are by default concurrent, so that statements that do not share common data can run concurrently.


    • defines functional computations separately from the data access logic, so that functional computation does not need to be respecified when a model compiler changes the data structures.


    • allows direct manipulation of UML elements only, so that model compilers can safely assume their own rules are not violated.



    The action semantics does not specify software structure anywhere. There are no mechanisms to denote persistence, or the manner of an invocation, or distribution, or how data is stored. All this is properly the business of the model compiler.


    An Executable UML model specifies the minimum required to show how a domain works in the context of the problem, and that's all.


    These requirements, and others, are discussed in detail in Software-Platform-Independent, Precise Action Specifications for UML [3].



    But this chapter is not about syntax. Accordingly, this chapter does not describe every syntactic element of the language we use (if statements and loops, for example), nor does it describe every syntactic nitty-gritty detail.




    To describe syntax, we use conventional syntax description forms, such as <class>, and we use the following typographical conventions for representing action language:



    Boldfaced words are keywords or reserved words.


    Capitalized Italics are class names.


    Lowercase italics are object references or attributes.



    For example:



    select any objectRef from instances of Class
    where selected.className == "my Class";
    // Comments may appear anywhere;
    // we show them below the action fragment



    To define a language is not difficult. Given the standard definition of the semantics that is provided by the action semantics, you could design your own favorite syntax. To illustrate this point, we present two other action languages in Section 7.6: Other Action Languages.










      I l@ve RuBoard



      10.8 Handling Duplicate Index Values




      I l@ve RuBoard










      10.8 Handling Duplicate Index Values




      10.8.1 Problem



      Your input contains
      records that duplicate the values of unique keys in existing table
      records.





      10.8.2 Solution



      Tell LOAD DATA to ignore the
      new records, or to replace the old ones.





      10.8.3 Discussion



      By default, an error occurs if you attempt to load a record that
      duplicates an existing record in the column or columns that form a
      PRIMARY KEY or
      UNIQUE index. To control this behavior,
      specify IGNORE or REPLACE after the
      filename to tell MySQL to either ignore duplicate records or to
      replace old records with the new ones.



      Suppose you periodically receive meteorological data about current
      weather conditions from various monitoring stations, and that you
      store measurements of various types from these stations in a table
      that looks like this:



      CREATE TABLE weatherdata
      (
      station INT UNSIGNED NOT NULL,
      type ENUM('precip','temp','cloudiness','humidity','barometer') NOT NULL,
      value FLOAT,
      UNIQUE (station, type)
      );


      To make sure that you have only one record for each station for each
      type of measurement, the table includes a unique key on the
      combination of station ID and measurement type. The table is intended
      to hold only current conditions, so when new measurements for a given
      station are loaded into the table, they should kick out the
      station's previous measurements. To accomplish this,
      use the REPLACE keyword:



      mysql> LOAD DATA LOCAL INFILE 'data.txt' REPLACE INTO TABLE weatherdata;









        I l@ve RuBoard



        Hack&nbsp;65.&nbsp;Detect Network Intruders with snort










        Hack 65. Detect Network Intruders with snort




        Let snort watch for network intruders and log attacksand alert you when problems arise.


        Security is a big deal in today's connected world. Every school and company of any decent size has an internal network and a web site, and they are often directly connected to the Internet. Many connected sites use dedicated firewall hardware to allow only certain types of access through certain network ports or from certain network sites, networks, and subnets. However, when you're traveling and using random Internet connections from hotels, cafes, or trade shows, you can't necessarily bank on the security that your academic or work environment traditionally provides. Your machine may actually be on the Net, and therefore a potential target for script kiddies and dedicated hackers anywhere. Similarly, if your school or business has machines that are directly on the Net with no intervening hardware, you may as well paint a big red bull's-eye on yourself.


        Most Linux distributions nowadays come with built-in firewalls based on the in-kernel packet-filtering rules that are supported by the most excellent iptables package. However, these can be complex even to iptables devotees, and they can also be irritating if you need to use standard old-school transfer and connectivity protocols such as TFTP or telnet, since these are often blocked by firewall rule sets. Unfortunately, this leads many people to disable the firewall rules, which is the conceptual equivalent of dropping your pants on the Internet. You're exposed!


        This hack explores the snort package, an open source software intrusion detection system (IDS) that monitors incoming network requests to your system, alerts you to activity that appears to be spurious, and captures an evidence trail. While there are a number of other popular open source packages that help you detect and react to network intruders, none is as powerful, flexible, and actively supported as snort.



        7.4.1. Installing snort


        The source code for snort is freely available from its home page at http://www.snort.org. At the time this book was written, the current version was 2.4. Because snort needs to be able to capture and interpret raw Ethernet packets, it requires that you have the Packet Capture library and headers (libpcap) installed on your system. libpcap is installed as a part of most modern Linux distributions, but it is also available in source form from http://www.tcpdump.org.


        You can configure and build snort with the standard configuration, build, and install commands used by any software package that uses autoconf:



        $ tar zxf snort-2.4.0.tar.gz
        $ cd snort-2.4.0
        $ ./configure
        [much output removed]
        $ make
        [much output removed]



        As with most open source software, installing into /usr/local is the default. You can change this behavior by specifying a new location, using the configure command's --prefix option. To install snort, su to root or use sudo to install the software to the appropriate subdirectories of /usr/local using the standard make install command:



        # make install



        At this point, you can begin using snort in various simple packet capture modes, but to take advantage of its full capabilities, you'll want to create a snort configuration file and install a number of default rule sets, as explained in the next section.




        7.4.2. Configuring snort


        snort is a highly customizable IDS that is driven by a combination of configuration statements and loadable rule sets. The default snort configuration file is the file /etc/snort.conf, though you can use a configuration file in any location by specifying the full path to and name of the configuration file using the snort command's -c option. The snort source package includes a generic configuration file that is preconfigured to load many sets of rules, which are also available from the snort web site at http://www.snort.org/pub-bin/downloads.cgi.



        To get up-to-the-minute rule sets, subscribe to the latest snort updates from the SourceFire folks, the people who wrote, support, and update snort. Subscriptions are explained at http://www.snort.org/rules/why_subscribe.html. This is generally a good idea, especially if you're using snort in a business environment, but this hack focuses on using the free rule sets that are also available from the snort site.




        It's perfectly fine to create your own configuration file, but since the template provided with the snort source is quite complete and shows how to take advantage of many of the capabilities of snort, we'll focus on adapting the template configuration file to your system.


        To begin customizing snort, su to root and create two directories that we'll use to hold information produced by and about snort:



        # mkdir -p /var/log/snort
        # mkdir -p /etc/snort/rules



        The /var/log/snort directory is required by snort; this is where alerts are recorded and packet captures are archived. The /etc/snort directory and its subdirectories are where I like to centralize snort configuration information and rules. You can select any location that you want, but the instructions in this hack will assume that you're putting everything in /etc/snort.


        Next, cd to /etc/snort and copy the files snort.conf and unicode.map to the parent directory (/etc). The /etc directory is the default location specified in the source code for these core snort configuration files. As we'll see in the rest of this hack, we'll put everything else in our own /etc/snort directory.


        Now you can bring up the file /etc/snort.conf in your favorite text editor (which should be emacs, by the way), and start making changes.


        First, set the value of the HOME_NET variable to the base value of your home or business network. This prevents snort from logging outbound and generic intermachine communication on your network unless it triggers an IDS rule.



        If the machine on which you'll be running snort gets its IP address via DHCP, you can set HOME_NET using the declaration var HOME_NET $eth0_ADDRESS, which sets the variable to the IP address assigned to your Ethernet interface. Note that this will require restarting snort if the interface goes down and comes back up while snort is running.




        Next, set the variable EXTERNAL_NET to identify the hosts/networks from which you want to monitor traffic. To avoid logging local traffic between hosts on the network, the most convenient setting is !$HOME_NET:



        var EXTERNAL_NET !$HOME_NET




        Forgetting the $ is a common mistake that will generate an error about snort not being able to resolve the address HOME_NET. Make sure you include the $ so that snort references the value of the $HOME_NET variable, not the string HOME_NET.




        If your network runs various servers, the next step is to update the configuration file to identify the hosts on which they are running. This enables snort to focus on looking for certain types of attacks on systems that are actually running those services. snort provides a number of variables for various services, all of which are set to the value of the HOME_NET variable by default:



        # List of DNS servers on your network
        var DNS_SERVERS $HOME_NET
        # List of SMTP servers on your network
        var SMTP_SERVERS $HOME_NET
        # List of web servers on your network
        var HTTP_SERVERS $HOME_NET
        # List of sql servers on your network
        var SQL_SERVERS $HOME_NET
        # List of telnet servers on your network
        var TELNET_SERVERS $HOME_NET
        # List of snmp servers on your network
        var SNMP_SERVERS $HOME_NET



        Next, copy the classification.config and reference.config files to /etc/snort and set the include statements for these in snort.conf to point to the full path to these files:



        include /etc/snort/classification.config
        include /etc/snort/reference.config



        Now set the value of the RULE_PATH variable in the snort configuration file to /etc/snort/rules (this variable can point anywhere, of course, but I prefer to centralize as much of the snort configuration information in /etc/snort as possible):



        var RULE_PATH /etc/snort/rules



        Finally, configure snort's output plug-ins to log rule transgressions (known as alerts) however you'd like. By default, snort enables you to log alerts to the system log and various databases, and also makes it easy for you to define custom alert mechanisms. I'll focus on using the system log, since that's the most common (and generic) logging mechanism. To enable logging alerts to the system log (/var/log/messages), simply uncomment the following line in /etc/snort.conf:



        output alert_syslog: LOG_AUTH LOG_ALERT



        Almost there! You're now ready to download and install the rules files that are referenced in your snort configuration file. As mentioned previously, you should seriously consider subscribing to these if you're using snort in an enterprise environment, both in order to support further development of snort and because it's simply the right thing to do. For the purposes of this hack, you can retrieve and install the free (unregistered user) rules files from http://www.snort.org/pub-bin/downloads.cgi by searching the page for the "unregistered user release" section and retrieving a gzipped tarball of the rules that match the version of snort you've built.


        To install these rules, change directory to your /etc/snort directory and su to root or use sudo to extract the contents of the tarball with a standard tar incantation:



        $ cd /etc/snort
        $ sudo tar zxvf /home/wvh/snortrules-pr-2.4.tar.gz



        This will create /rules and /doc subdirectories in /etc/snort. (Again, these rules can actually live anywhere on your system since their location is identified by the RULE_PATH variable in the snort configuration file. We set this variable to /etc/snort/rules earlier.)




        7.4.3. Starting snort



        At this point, you're ready to run snort. Though snort offers a daemon mode, it's generally useful to run it in interactive mode from the command line until you're sure you've made the correct modifications to your /etc/snort.conf file. To do this, execute the following command:



        # snort -A full



        You'll see a lot of output as snort parses your configuration file and rule sets. If you've done everything right and not made any typos, this output will conclude with the following block of output:



        --== Initialization Complete ==--

        ,,_ -*> Snort! <*-
        o" )~ Version 2.4.0 (Build 18) x86_64
        '''' By Martin Roesch & The Snort Team: http://www.snort.org/team.html
        (C) Copyright 1998-2005 Sourcefire Inc., et al.



        If you see this, all is well and snort is running correctly. If not, correct the problems identified by the snort error messages (which are usually quite good), and try the snort command again until snort starts correctly.


        One especially common and irritating message when getting started using snort is the following:



        socket: Address family not supported by protocol



        You will see this message if your system's kernel is not configured to support the CONFIG_PACKET option, which enables applications (the packet capture library, in this case) to read directly from network interfaces. This capability can be compiled directly into the kernel, but it's more commonly built as a loadable kernel module (LKM) with the name af_packet.ko (af_packet.o if you're still running a pre-2.6 Linux kernel).


        If this capability is provided as an LKM on your system, you can generally load it by executing the modprobe af_packet.ko command as root or via sudo. If modprobe doesn't work for some reason, you can load the module directly using the insmod command. The name of the appropriate /lib/modules subdirectory where the module is located is contingent on the version of the kernel you're running, which you can determine by executing the uname -r command. For example:



        # uname -r
        2.6.11.4-21.8-default
        # insmod /lib/modules/2.6.11.4-21.8-default/kernel/net/packet/af_packet.ko
        Testing Snort



        The fact that snort is running without complaints is all well and good, but executing correctly isn't the same thing as doing what you want it to do. It's therefore useful to actually test snort by triggering one of its rules. The easiest of these to trigger are the port scan rules. To test these, connect to a machine outside your network and issue the nmap command, identifying the machine on which you're running snort as the target, as in the following example:



        $ nmap -P0 24.3.53.235
        Starting nmap V. 2.54BETA31 ( www.insecure.org/nmap/ )
        Warning: You are not root -- using TCP pingscan rather than ICMP
        Nmap run completed -- 1 IP address (0 hosts up) scanned in 60 seconds



        You can now check /var/log/snort, in which you should see a filenames alert with contents like the following:



        a[**] [122:17:0] (portscan) UDP Portscan [**]
        09/14-20:53:16.024463 24.3.53.235 -> 192.168.6.64
        RAW TTL:0 TOS:0xC0 ID:29863 IpLen:20 DgmLen:163



        You will also see a directory with the name 24.3.53.235. This directory contains logs of the offending packets that triggered the alert. Congratulations! snort is working correctly.



        If you have port forwarding active on a home or business gateway, you'll probably see a file with the IP address of the gateway instead of the IP address of the host from which you did the port scan.




        Once you're satisfied that snort is working correctly, you'll probably want to terminate the interactive snort session we started earlier and restart snort in daemon mode, using the following command:



        # snort -A full -D



        This starts snort in the background and sends its initialization messages to /var/log/messages. To add this command to your system's startup mechanisms, either append it to a startup script such as /etc/rc.local or integrate it into the standard system startup process by creating a start/stop script in /etc/init.d and adding the appropriate symbolic links to the /etc/rc.runlevel directory that corresponds to the default runlevel for the system on which you're running snort.




        7.4.4. Advanced snort


        You can extend snort in an infinite number of ways. One of the easiest is to take advantage of more of its default capabilities by activating additional rule sets that are provided in the bundle that you downloaded but are commented out of the default snort configuration file template. Some of my favorites to uncomment are the following:



        include $RULE_PATH/web-attacks.rules
        include $RULE_PATH/backdoor.rules
        include $RULE_PATH/shellcode.rules
        include $RULE_PATH/virus.rules



        Once you uncomment these and restart snort, you'll probably start to see additional snort alerts such as the following:



        [**] [1:651:8] SHELLCODE x86 stealth NOOP [**]
        [Classification: Executable code was detected] [Priority: 1]
        09/15-04:49:32.299135 70.48.80.189:6881 -> 192.168.6.64:52757
        TCP TTL:109 TOS:0x0 ID:53803 IpLen:20 DgmLen:1432 DF
        ***AP*** Seq: 0x1869E9D1 Ack: 0x18F60ED8 Win: 0xFFFF TcpLen: 32
        TCP Options (3) => NOP NOP TS: 719694 594700245
        [Xref => http://www.whitehats.com/info/IDS291]



        Better to know about attempted attacks than to be blissfully unaware! Of course, whether or not you want to monitor your network for these types of attacks is entirely dependent on your site's network policieswhich is why they're commented out of the snort configuration file template. Your mileage may vary, but I find these quite useful.




        7.4.5. Summary


        snort is an extremely powerful, flexible, and configurable intrusion detection system. This hack focused on getting it up and running in a standard fashionexplaining how to create your own rules and take advantage of all of its capabilities would require its own book. Actually, a number of books on snort are available, as well as extensive discussions in more general networking texts such as O'Reilly's own Network Security Hacks, by Andrew Lockhart.


        If you're interested in a simpler network-monitoring package, PortSentry (http://sourceforge.net/projects/sentrytools/) is one of the best known, though it hasn't been updated for quite a while now. However, snort is a much more powerful tool and is actively under development. Newer snort developments include the ability to actively respond to certain types of attacks by sending certain types of packages (known as flexresp, or flexible response) and increasing integration with dynamic notification tools on both the Linux and Windows platforms. In today's connected world, you can't really afford not to firewall your hosts and scan for clever folks that can still punch through your defenses. In the open source world, there's no better tool for the latter task than snort.




        7.4.6. See Also


        • "Monitor Network Traffic with MRTG" [Hack #79]

        • Network Security Hacks, by Andrew Lockhart (O'Reilly)

        • man snort

        • Snort Central: http://www.snort.org